Pith. sign in

Paper Citation Record · LEDGER

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation

As of 23 August 2026, this Paper Citation Record lists 100 of 140 outbound references and 0 inbound Pith citation observations for arXiv:2606.07936.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.07936 v2

Coverage vector

measured 100 of 140 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T20:20:08.996005Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 140 outbound references displayed

  • verified exact26
  • verified fuzzy0
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch11

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b16c6cd-27a3-415d-94c2-2b0f64883f7d · outbound

This paper cites Scientific reports , volume=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Scientific reports , volume=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:9ed42ab8aa9d4c1a6d28596dbf4980036eb30bc8151a9d3dc3e664e9a2d48a06

Observation 98912337-64b8-45d6-be81-e4a388d28567 · outbound

This paper cites A Critical Evaluation of Evaluations for Long-form Question Answering.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation A Critical Evaluation of Evaluations for Long-form Question Answering

Reference 2

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.088220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:a16778d8c72199036ee1fc71ebb4ce663bb42cad6db9334a461b8209bf761579

Observation 1b430ddb-1507-4fcc-b52a-c5d9be28eed1 · outbound

This paper cites Responsible AI Considerations in Text Summarization Research: A Review of Current Practices.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Responsible AI Considerations in Text Summarization Research: A Review of Current Practices

Reference 3

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.090211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:171c76f1ad53e0cafd9112bf24104d4fcb811946303b55c0bfe66237fb0adf52

Observation 2333c820-7649-4aab-9860-e5bd0531c817 · outbound

This paper cites Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations

Reference 4

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.092003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:f526bae6b50a83fccb292e7d8626f78588b80e4345cdda13f90ef7465aff2c41

Observation 63ca7633-8eed-4921-bf5e-67a5c7b7bb23 · outbound

This paper cites 2025 , address=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation 2025 , address=

Reference 5

Resolution
metadata mismatch
doi, observed 2026-06-27T20:21:13.086322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:5bf2fbd52b21e270853886f8904189201fc8f43844552effdf3f731ed9d51776

Observation 7b62a59c-d62b-40a8-950f-a8e86f33fbef · outbound

This paper cites First Conference on Language Modeling , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation First Conference on Language Modeling , year=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:cb1cf3b7eb7eef004d640ad9742532bc4b2f475d65904a18abc5deb87ad49032

Observation 9f402b79-42c9-4c75-bb00-c16d1a5cf483 · outbound

This paper cites Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM ^2 ) , pages=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM ^2 ) , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:7bd4b929c961e1fd66115698ec76d31d3446a91dd553115c9c4f365c2a66f969

Observation 92e8be16-1bd9-4e99-8479-f931ee0db13d · outbound

This paper cites NPJ digital medicine , volume=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation NPJ digital medicine , volume=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:be108e998019503abc0cd64f8a7701f509938a178de314d5cd4a4c714d5a62a8

Observation 7851f010-c71e-4784-819d-5faf2b95d91b · outbound

This paper cites Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts , pages=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts , pages=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:84223b1e22e3404973d103b32d4252b41e3ca09aff6f3bb64db09b7942cca886

Observation 0adf9169-21be-4635-b4f3-e681ab698b64 · outbound

This paper cites Proceedings of the 2nd Workshop on Human Evaluation of NLP Systems (HumEval) , pages=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 2nd Workshop on Human Evaluation of NLP Systems (HumEval) , pages=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:b1c631b17e8fe7c0a0009c643b56327e89775df3c169ecd51f9f3ea5f507fe0f

Observation 0281dc86-eae2-4c2d-9e48-d6d8d5c9fb87 · outbound

This paper cites Proceedings of the Fourth Workshop on Insights from Negative Results in NLP , pages=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the Fourth Workshop on Insights from Negative Results in NLP , pages=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:be166925123a07f36e755eb0e536af1d45bbadf09c3ec5b798488fa8f2b2dc0a

Observation ef034d17-1977-43df-b007-e9d0303743bb · outbound

This paper cites Computational Linguistics , volume=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Computational Linguistics , volume=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:87882f0ad4f4f5974e2b44edbe29d938fcba0b8b18bdf917313951f1ceab059b

Observation af3084de-54ff-4635-ab0c-ace829bba0a9 · outbound

This paper cites University of Chicago Coase-Sandor Institute for Law & Economics Research Paper , number=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation University of Chicago Coase-Sandor Institute for Law & Economics Research Paper , number=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:eb6c91eb44fd5be93c163e7bc67ff04dd9709999f920af9f88ad25ee4f052bf3

Observation fdea24ff-4e05-4d6a-8772-a5ece13379bd · outbound

This paper cites Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM ^2 ) , pages=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM ^2 ) , pages=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:1cfd6cc7e87b816a5868daae84b714a043f2089799f6638847e283aa2f25b510

Observation 88298a41-01d8-42aa-8d7b-64f7657b6375 · outbound

This paper cites LLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation LLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks

Reference 15

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.082428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:668e972162c042ee2cf46fd0a50bfbff6ece64f3bb574c61a9250292baac82fb

Observation 4d438264-bbc2-4a75-afce-c3cb7c8db1cf · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Advances in Neural Information Processing Systems , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:4a29f92eae11c3a905803dbe5625ae515059b92639df46244acb696d46e21cc6

Observation b7354b67-6b30-476d-8917-49ef0f9f56ea · outbound

This paper cites npj Health Systems , volume=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation npj Health Systems , volume=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:896911ebb613d83f279492bdddd70183a47e9f1a6d423a68cb9afd8c76db943a

Observation 95367f82-0974-4762-8d4d-4035c3b2d267 · outbound

This paper cites Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:6003e57c7bb423832a1a8a3647a8427fcc51fb5ce339bf238d7faf137c22c925

Observation 04d912d3-043c-4af2-b405-f0741be2e3df · outbound

This paper cites Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:2f1fe0e42dd1d005719c0773b956864877fb84bb046b4e0b99258b6bd4116ec4

Observation bcf63442-4f57-444a-9a5a-b5e8ec1ef364 · outbound

This paper cites Proceedings of the 31st International Conference on Computational Linguistics , pages=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 31st International Conference on Computational Linguistics , pages=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:c1baf7d3e49f51085598b1cd31c9a37fa5aa1bdab31d3f1542b58e15fdfa4110

Observation dfe87851-4caa-42ba-b253-e3910c92ec55 · outbound

This paper cites The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:37:22.658232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:f5000e207bc6af97d23310523017794ee8ecbfda8f224e3ecdf283c3f040a34e

Observation cd910f03-e810-4dc0-8a36-b4b13e7c6770 · outbound

This paper cites Nature human behaviour , volume=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Nature human behaviour , volume=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:4b8b5250dcbfedc6cfb6cfd0238fec61ad96847b4bc069dfc55efae771e60171

Observation 31fbb181-f169-4e58-bc82-458b9d881756 · outbound

This paper cites 2022 , url =.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation 2022 , url =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:0269d18eb802242835e8ebedfb1a97a210182f9c8928410937c1b0170fb8f618

Observation 01f1039a-74f2-43d3-a22e-a95ece5e03fa · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 24

Resolution
metadata mismatch
doi, observed 2026-06-27T20:21:13.077833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:b6ac377c7fa609294e08bc607d0b521bbdd26aaff4cba4dd8d8a12a444855b9c

Observation b523a3cc-2551-4327-bd6d-f8f1b0376a0f · outbound

This paper cites ArXiv , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:f92b382ec18a845c23d527c0e141cf9d693e170798ab4681d607473c712ed6e0

Observation 2bd170a5-8424-4fbd-9a61-b7a86e51a56e · outbound

This paper cites PaperQA: Retrieval-Augmented Generative Agent for Scientific Research.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation PaperQA: Retrieval-Augmented Generative Agent for Scientific Research

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:37:22.670917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:59830463794ceeb6bb5b28d7f63aba6d8fc679963a3ba4515ac8fbb75dbd5164

Observation 9c9ebfc8-6d43-4280-bf46-7259751c6732 · outbound

This paper cites arXiv e-prints , pages=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation arXiv e-prints , pages=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:a4b01f82d79fbd79cf35051cb7ed44599f9b2388a9baf0256d9bb995b0a1fa07

Observation ca1cb1e1-32dc-436e-97cc-8d24ac2cce54 · outbound

This paper cites OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:37:22.686246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:2bc5ee820e23f0e916cd8083284356b33c41bbea6d29d5a0d8e0c41041aac1f1

Observation 4ade0dea-bbf2-4462-b04d-0f42eedda8df · outbound

This paper cites Text summarization branches out , pages=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Text summarization branches out , pages=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:5d28e7ba391c263096ec79fa8d424ebd6377acb279818e10d7b35f617ba75508

Observation b5a05ed5-8ba4-4a53-997e-4f5d9c57bedf · outbound

This paper cites Mistral 7B.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Mistral 7B

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T20:37:22.694387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:48d98a62e37ea8313cd585e1516ad7afc62581d8e0d6459fe64e5b312d28bcd1

Observation 17073da2-dec6-4adf-ac80-1ebfb87ec695 · outbound

This paper cites GPT-4 Technical Report.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation GPT-4 Technical Report

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T20:37:22.673224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:0e02398fe0cdc73d0c1086d77b470c5e98d62ed0df187818fc099d7f6364b99a

Observation 3c2823db-2c19-48de-a9f6-3a5156808422 · outbound

This paper cites Clinical Natural Language Processing Workshop , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Clinical Natural Language Processing Workshop , year=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:b80e589deb8e192e75149acc81f50b10e442f8b15f922f941baf4795078a69d4

Observation d7d9876d-33c0-44ff-b2e2-f083aac8379b · outbound

This paper cites Conference on Empirical Methods in Natural Language Processing , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Conference on Empirical Methods in Natural Language Processing , year=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:ab437095b1c70cd6b03ed18f72499410502c45e581efc48e834f6477559479db

Observation 0c7ddf9c-ab5e-44ac-b4e5-623e0f1957c7 · outbound

This paper cites Annual Meeting of the Association for Computational Linguistics , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Annual Meeting of the Association for Computational Linguistics , year=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:14c90273f41d761c1c11a59b14c9180c05007d2eb0004994628e969e3c91e3ca

Observation 7a40d2b5-b97e-49b3-900d-768dd06ffe68 · outbound

This paper cites Annual Meeting of the Association for Computational Linguistics , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Annual Meeting of the Association for Computational Linguistics , year=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:5f58ca2554a8aebe804a8124930dcff62e5a35900c1cb311c2576255e2cd26f5

Observation d86d1e9b-946b-4b4c-9040-e7ceefc304af · outbound

This paper cites Annual Meeting of the Association for Computational Linguistics , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Annual Meeting of the Association for Computational Linguistics , year=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:3cc2b2c74ba8d85a27ea9e35ae4650d3293b2b0b7295a0a20674a5b81695a32f

Observation dfb158c3-27fc-47ef-80ac-c778966136db · outbound

This paper cites ArXiv , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:f98890398f2393bdd668e0b4f405306110ef9d08431413d1648d4614d140140f

Observation 989d5e63-135a-41d1-8d38-37dea9f1b515 · outbound

This paper cites , author=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation , author=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:68af6836f0ebd2b554d6007daab956262f42eaeee2a063d483b80e5f2159070f

Observation 680c6a2d-57a6-4bd9-8e20-37e1514d6030 · outbound

This paper cites ArXiv , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:0117bfb0cdbd54b2eb24add60170cf9ff781c24e5f1cf9679ade5c326fa37917

Observation 01d7754b-ca66-4846-ab93-90daf0fd7095 · outbound

This paper cites , author=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation , author=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:b1a9c8519f72715490702d615fcea12ecef3296791212812493f24223b3eaf8a

Observation d82d8d15-5815-49dc-9a3e-dcef00ea0bd7 · outbound

This paper cites , author=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation , author=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:47b4098055e9607132caeea43d973b1db72465ae5ec9ea9dcc6a27f5ab8fb650

Observation 79e4abbe-3665-4ae9-b218-ec7115280f91 · outbound

This paper cites Yu, Qiang Yang, and Xing Xie.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Yu, Qiang Yang, and Xing Xie

Reference 42

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.063575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:e41f08cf997955f77db81efc870a9da6f45cede801423601d6cc53df7e6abb02

Observation 79981ea4-a605-417e-bc27-999a4909d88e · outbound

This paper cites Improving Factuality in Clinical Abstractive Multi-Document Summarization by Guided Continued Pre-training.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Improving Factuality in Clinical Abstractive Multi-Document Summarization by Guided Continued Pre-training

Reference 43

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.050279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:636568384162b941a5b8402b3945004347f8c4e93e357c672638f9161cf11774

Observation 3d4ed799-7988-4fa4-af97-67c11843529c · outbound

This paper cites FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:37:22.675805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:17f789f5ad54bd58b5376281e45692d89859b6c04f8a37f3d9d2aa0fc4455ebe

Observation 50e44a69-99ad-4685-886a-608f25a6e51c · outbound

This paper cites Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:37:22.689103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:a037d0ff6dfc580358e0085e967748b1a5e7205dd9d31bec52fce180e637f349

Observation 42a97c9f-9026-4ebe-aa8f-a0d5c1ea47b9 · outbound

This paper cites Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:ac69ab84c9a385ba690bc0d1ed45eefe8ee0838633ebf52edf48f9aa1912d671

Observation 8578f1d1-db02-472f-b8a9-4aa23a7a097a · outbound

This paper cites medRxiv , pages=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation medRxiv , pages=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:165440b996bfbccb156daa32c00748d3136c37b0e4555be12ee4b7bffd79c67d

Observation d4fa0a0c-3c0f-48ee-9634-649bcf972adc · outbound

This paper cites and Villarroel, Mauricio and Clifford, Gari D.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation and Villarroel, Mauricio and Clifford, Gari D

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:7ac84b2792a263869d40a13dce2aa794ded8493eec197e9ba97eaf4360161ffe

Observation a75cdf03-cbb2-47ce-a62c-5c715c5b1507 · outbound

This paper cites ArXiv , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:71f76cbddde8807059f94c1e334f2e7b05c965dc106a2875121d3e80d50d7554

Observation 6910d7ec-6ce3-40ad-9f75-7ac69f68de76 · outbound

This paper cites Bioinformatics , volume =.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Bioinformatics , volume =

Reference 50

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.065937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:509e4f412893ad98c14a194fea93bc62ca713287f38fff32c0742370c4aee866

Observation e77617e3-a584-442e-8909-e1c024736c16 · outbound

This paper cites TOPICAL : TOPIC Pages A utomagica L ly.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation TOPICAL : TOPIC Pages A utomagica L ly

Reference 51

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.059575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:837f933da89b0f57771be35288a8b66acb8bd88c828f71d63d04e888df5e8dea

Observation 500462cd-5edb-4a08-a252-2a017dfee225 · outbound

This paper cites What’s in a Summary? Laying the Groundwork for Advances in Hospital-Course Summarization.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation What’s in a Summary? Laying the Groundwork for Advances in Hospital-Course Summarization

Reference 52

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.073952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:3b5091683b323927ef6a1dba110074ce3e78f1979dd371808a17d1eef4d94bb3

Observation d629c238-c82e-42a1-83b7-272d3e0ee3b7 · outbound

This paper cites 2015 26th international workshop on database and expert systems applications (dexa) , pages=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation 2015 26th international workshop on database and expert systems applications (dexa) , pages=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:8cd3ccaba92f863489b2177f7a8a475134b93d2091ad717d930d17a464b8a072

Observation 096afb44-0e96-4edd-b81d-1725cbfabe3c · outbound

This paper cites Journal of Intelligent Connectivity and Emerging Technologies , volume=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Journal of Intelligent Connectivity and Emerging Technologies , volume=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:e9a20e474a5bb1bc067c93adf9c87f2f2dfa4bebcb3e4f77c52c41649edbc5f2

Observation 6c57ae8e-dbbc-4b85-bf9e-d929a1b47e12 · outbound

This paper cites Journal of biomedical informatics , volume=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Journal of biomedical informatics , volume=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:1bbbbb80325b2af544af579f1dd1e7ce4bdd14cea1dd288c98ad2419d58d648c

Observation 7222757e-d931-478a-b6d4-35362ae28b33 · outbound

This paper cites A Novel System for Extractive Clinical Note Summarization using EHR Data.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation A Novel System for Extractive Clinical Note Summarization using EHR Data

Reference 56

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.042771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:06716bd80fb4c8eb16d888c048612818bdffb7f5a0ca900bb54032992a34c631

Observation 93100199-bc46-4fef-aed2-dcc1c5576915 · outbound

This paper cites & Lipton, Z.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation & Lipton, Z

Reference 57

Resolution
metadata mismatch
doi, observed 2026-06-27T20:21:13.037312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:54ac34c348adb83a7b17b3925c4ee22db682e613048d9284acd655082afc6a7d

Observation 02a5fc08-22d1-44df-bc55-00b3c8ab39dc · outbound

This paper cites Towards Automating Medical Scribing : Clinic Visit D ialogue2 N ote Sentence Alignment and Snippet Summarization.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Towards Automating Medical Scribing : Clinic Visit D ialogue2 N ote Sentence Alignment and Snippet Summarization

Reference 58

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.040837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:87451c036801cb84683733caf0f6d065945d3b17342dcf710b385c2002e72911

Observation 35044527-a97a-4540-b321-b02590d9c47a · outbound

This paper cites DERA : Enhancing Large Language Model Completions with Dialog-Enabled Resolving Agents.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation DERA : Enhancing Large Language Model Completions with Dialog-Enabled Resolving Agents

Reference 59

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.076027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:ffb6312ad43d1dc1681ba5243ba06e2ac2ffff1585aa8d96c040875f083993a9

Observation 66a7e94b-90ac-4811-b637-e256c3f00507 · outbound

This paper cites JMIR medical education , volume=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation JMIR medical education , volume=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:33f5f350638256bf7645a8718931e483b52265d20eba0812b4ddb2d6fcb865c2

Observation 3858b38e-2fb3-4e28-b94a-03ee0238e5df · outbound

This paper cites Cureus , volume=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Cureus , volume=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:a8906ce7978a8e7bf6f792e3cbad7790a6c5e4884dc45522419d8c3e7577c3aa

Observation f1d50d2e-fa1b-4217-9241-dbd2f66e2458 · outbound

This paper cites ArXiv , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:11735e0ca2ec0da1c29e9817d4de4d7e8e8ab660bb641572fadb4824a0838300

Observation 83ced66a-83d2-456d-86b9-966a6712b6c2 · outbound

This paper cites From Protocol to Screening: A Hybrid Learning Approach for Technology-Assisted Systematic Literature Reviews.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation From Protocol to Screening: A Hybrid Learning Approach for Technology-Assisted Systematic Literature Reviews

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:37:22.694249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:294f366ae171dd434c6d87c02007137dd221462c2e20d80c262448092a1b7a18

Observation 499f5a0b-782a-4ff5-bc47-dfb225598d0a · outbound

This paper cites Hashimoto.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Hashimoto

Reference 64

Resolution
metadata mismatch
doi, observed 2026-06-27T20:21:13.034125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:52eb56f526492d206ee483b9356c41c9974acee728992a67a572ea95acd988e6

Observation 3d50d6e7-656c-4b2d-b1d7-6fe1267d4be4 · outbound

This paper cites On Learning to Summarize with Large Language Models as References.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation On Learning to Summarize with Large Language Models as References

Reference 65

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.080392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:1278cb872727967d861c635cba408269758e78fb49b0280f84d6f38cd47992aa

Observation d009cb8c-f181-432d-a6d7-9cfca7c6bfc0 · outbound

This paper cites Summarizing, simplifying, and synthesizing medical evidence using GPT -3 (with varying success).

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Summarizing, simplifying, and synthesizing medical evidence using GPT -3 (with varying success)

Reference 66

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.056076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:f9ba692eba9efbe120fde2303b44fe64ca2feb5de60d5399a4cc0aca31b488f6

Observation d4d46906-9517-447a-802f-13ff88d86416 · outbound

This paper cites An empirical survey on long document summarization: Datasets, models, and metrics.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation An empirical survey on long document summarization: Datasets, models, and metrics

Reference 67

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.048508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:caa07690b3d8556ff6453bca6cdb16d2a9401deddf81043abe1ac64c287f58b2

Observation f4033730-755b-4562-918d-2a7897fc2b42 · outbound

This paper cites Conference on Empirical Methods in Natural Language Processing , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Conference on Empirical Methods in Natural Language Processing , year=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:78c244fd9125c8023f778356f66a7c3a87de1074613a6f3635b26dca666e9587

Observation cf87f024-4abc-49d1-928d-8dc81c87a1e7 · outbound

This paper cites ACM Trans.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ACM Trans

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-06-27T20:21:13.073698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:e5a77e8b4d0dffbfcc0773000398f85d701a4a56fd0cf458fa0102ca7e2a804e

Observation c59feabe-6e95-4cd8-955d-553e116ae950 · outbound

This paper cites 2023 , journal =.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation 2023 , journal =

Reference 70

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.046915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:332ca6d8b6751bc9094cb9afe9d3f33cf1139a27e0a1cfb9b919cdaef1aa37f8

Observation 3525dbf9-67e6-4d0f-9ca2-c17c250cca5c · outbound

This paper cites Annual Meeting of the Association for Computational Linguistics , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Annual Meeting of the Association for Computational Linguistics , year=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:6e404e5077005d0bd953c7f54d12447995656e511642b5df72a451323107a71e

Observation 058447c0-3117-4e47-bf33-1e40dbf769ef · outbound

This paper cites ArXiv , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:718a5de27e23d9de624e707b8a8cabebac7419159e73df34d73efa22e02ca571

Observation df9f31c7-484f-4912-969d-a97d9286169e · outbound

This paper cites Journal of Medical Internet Research , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Journal of Medical Internet Research , year=

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:be13738b388dd9d6190a4e9dc4f5c7018b98706a3b03acffb2a8c0f032ae9121

Observation 23a06557-3ed7-47ad-8984-5345d7388e81 · outbound

This paper cites BMJ Open , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation BMJ Open , year=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:6164f763f1cf4eca8cdbda06b0dcefd6536ab2f7ccd6ac33332e5b9c29728fb0

Observation e273582a-ea3e-4508-98b3-f36c2d1b034b · outbound

This paper cites Energies , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Energies , year=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:45afb7e207cfb94339c3c222e8541edabec0f154c60aa8bfe5fd23d8cad5d0cd

Observation 71b6e608-f57d-4808-904c-a7a0a6935368 · outbound

This paper cites Proceedings of the 52nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 52nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:eef3ff9ec91486e472c32e9f92dd9a7001168e2fa99d07b494bc6ad46aa27e32

Observation 29700864-7123-40c8-a28a-0287941ffb0e · outbound

This paper cites an unresolved cited work.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:61b100b0b14bc5fb34aff366b5ec2072a63b22f7268331aba549bc904bb1ec94

Observation 2fdf8532-0405-4246-97b5-f7d5d0c3c282 · outbound

This paper cites Proceedings of the conference.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the conference

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:2ae4badf1ba728fb615f94c7d4b499fd92d8971d3b4961ee0f298986cc4c5666

Observation a0844e4b-fb9c-4de3-a675-905732047568 · outbound

This paper cites Trends in cognitive sciences , volume=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Trends in cognitive sciences , volume=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:d0529f359300348a83e256c6d7efc62bb4a9c524b70b919e748460f20f20edca

Observation 8a318ff8-fdff-459b-aea2-26b4ea2b1a21 · outbound

This paper cites Craik , abstract =.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Craik , abstract =

Reference 80

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.081689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:e90d7968c2a58fe9464a7e9e42da59e4dce6b5b75bcd16e73df65aa1f9feebff

Observation bdf59089-fc1d-4dc1-861a-6b7b06fcdbe2 · outbound

This paper cites Conference on Empirical Methods in Natural Language Processing , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Conference on Empirical Methods in Natural Language Processing , year=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:b218d2313a607ac83a754e026bf627abbdc2f733602299cbea0d48d07ec9e763

Observation add8bb77-99b8-44f8-a243-419ffef34766 · outbound

This paper cites ArXiv , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:ee1272aac41f7490c0e43263561ca15bd09ea7d81d310f3d377e3123e08c17a5

Observation b8041b7b-47db-417a-8afc-c4eb9c13c160 · outbound

This paper cites ArXiv , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:9e44a63f62c59e1f07e2a023c21905cbbd70307191baeece61c3d9fc427ce34c

Observation 40afc49a-429e-48b5-a658-cfab1b4c2c3e · outbound

This paper cites Annual Meeting of the Association for Computational Linguistics , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Annual Meeting of the Association for Computational Linguistics , year=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:4b5dd5defc26fa62b54f13cf34393db0ed9c4cf86bf91f1dfffce7236149f898

Observation 19d1e942-015b-4bf7-8162-42ab26d79b24 · outbound

This paper cites Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , year=

Reference 85

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:9826d3ed38a989b0eac82df1dd5767af060749fe8fab759ab4dea6e6e27cc49f

Observation ec1320b5-868c-48f4-b46d-4da9355d15af · outbound

This paper cites ArXiv , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:58c4e534492478aa8e52b67100a96387e0acee0fc6b987c708ade11a0d5ec196

Observation 5d914574-9092-439e-9e21-c88e7c18bdad · outbound

This paper cites H i S truct+: Improving Extractive Text Summarization with Hierarchical Structure Information.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation H i S truct+: Improving Extractive Text Summarization with Hierarchical Structure Information

Reference 87

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.027023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:6b3c760d231a35ffeea8a195e6730ef4624321da9cfd5e29826e51f7dddd6c87

Observation 997af391-06bd-48f0-a7e6-ce3fda19e9c1 · outbound

This paper cites What factors might be causing the significant deviations in my circadian rhythm patterns over the past 30 days?.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation What factors might be causing the significant deviations in my circadian rhythm patterns over the past 30 days?

Reference 88

Resolution
metadata mismatch
doi, observed 2026-06-27T20:21:13.057471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:852d1575e03b9f8a9efcf900dc2af549629432816e2d7a3b646bc97205128699

Observation 2418d2f6-894b-4402-835b-8e4857e8bdbc · outbound

This paper cites and Belz, Anya and Du s ek, Ond r ej and Mille, Simon and van der Lee, Chris and Reiter, Ehud and Santhanam, Shivani and Thomson, Craig.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation and Belz, Anya and Du s ek, Ond r ej and Mille, Simon and van der Lee, Chris and Reiter, Ehud and Santhanam, Shivani and Thomson, Craig

Reference 89

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:1880dc0c54696f07218284510ef43196248e3c2d289552e004eef9e9153aeef8

Observation addd4d8d-530d-45cd-8296-1f1f828e1ab5 · outbound

This paper cites G-Eval : NLG Evaluation using GPT-4 with Better Human Alignment.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation G-Eval : NLG Evaluation using GPT-4 with Better Human Alignment

Reference 90

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:e7563f871df8469fae25334eec5e8e7f9df62088e33ba9a46f0c9969cf69eb8d

Observation 81dc610d-89e7-45e1-b7e4-cf622802c639 · outbound

This paper cites A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Governance Framework.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Governance Framework

Reference 91

Resolution
metadata mismatch
doi, observed 2026-06-27T20:21:13.080084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:6f395809b23a77d57e6c3e39672f173232f091130bfb4f9c4d843668156d8ed5

Observation b8eb6e69-c084-4c87-b914-610b0f4c24ba · outbound

This paper cites HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:37:22.686429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:dff8aa5a565f9954d0b0467fa6a6692b1351c69fb80626112268d48c5eb36c5d

Observation 2cabe6f7-8e1a-4cda-abb1-132bd67c0256 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2023 , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Findings of the Association for Computational Linguistics: EMNLP 2023 , year=

Reference 93

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:20388f0879a6760338cc76120c58e631b4344c4c35bece7bc03b116d9a622f05

Observation 1eac47b5-1fd2-427e-9246-ef72f0dda100 · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:647e156b4cb4f1beeafef94783319a92e333b8a04a019917f7d644324cb90795

Observation d52780b4-b71b-4926-aea5-9ddee4067efc · outbound

This paper cites Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers).

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)

Reference 95

Resolution
verified exact
doi, observed 2026-06-27T20:21:13.024730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:5d2123c0f6d19dc9935a461d2ed24803ee0310f8dbfea196e5ab1ba700f2398d

Observation be54c590-61bb-49ff-b72c-860402bd6037 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Transactions of the Association for Computational Linguistics , volume=

Reference 96

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:388ccdf9ddc4a3224e1de021471d7aef70add81a833f741c9220d8fef6befa4b

Observation cca0ce1c-c752-4e57-a8d3-76d7e35c5565 · outbound

This paper cites FFCI: A Framework for Interpretable Automatic Evaluation of Summarization.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation FFCI: A Framework for Interpretable Automatic Evaluation of Summarization

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:37:22.673684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:1d20ddf8f03a49d4175fbc4709f7a2c361fe3c56b9fad9c20c9665676573cb82

Observation e2964047-2545-45f2-aa8e-40f212b2f934 · outbound

This paper cites Journal of the Royal Society of Medicine , volume =.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Journal of the Royal Society of Medicine , volume =

Reference 98

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:05730606496f4a53b75888dc83d0ba9fd4a51fdb644b5414be5b763fda6fa7f5

Observation 556ded4f-87cc-4e5c-97f3-cf358ea1493c · outbound

This paper cites The significant cost of systematic reviews and meta-analyses: A call for greater involvement of machine learning to assess the promise of clinical trials , journal =.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation The significant cost of systematic reviews and meta-analyses: A call for greater involvement of machine learning to assess the promise of clinical trials , journal =

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-06-27T20:21:13.055005Z

Source-reported events for the cited work

correction dated 2019-09-12. Source: crossref record 10.1016/j.conctc.2019.100450->10.1016/j.conctc.2019.100443:correction, observed 2026-07-11T03:16:42.407043+00:00. This notice travels one citation hop only.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:3f6deba00d2fc4d9ddbc31af6e6e1c19623cfaf941cb6373b063a74d43da4aeb

Observation 7a04601b-ddd8-4b7d-9cd9-903a4cb5e19e · outbound

This paper cites npj Digital Medicine , volume=.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation npj Digital Medicine , volume=

Reference 100

Resolution
unresolved
no resolver link, observed 2026-06-27T20:20:08.996005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:dc277f58c1e686143a23c8052b514f291388781bd7fb2ba02ac8aae865433220

Pith citing papers

No inbound Pith citation observations are available.