Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T20:20:08.996005Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 100 of 140 outbound references and 0 inbound Pith citation observations for arXiv:2606.07936.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T20:20:08.996005Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 140 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3b16c6cd-27a3-415d-94c2-2b0f64883f7d · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Scientific reports , volume=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98912337-64b8-45d6-be81-e4a388d28567 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation A Critical Evaluation of Evaluations for Long-form Question Answering
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1b430ddb-1507-4fcc-b52a-c5d9be28eed1 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Responsible AI Considerations in Text Summarization Research: A Review of Current Practices
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2333c820-7649-4aab-9860-e5bd0531c817 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 63ca7633-8eed-4921-bf5e-67a5c7b7bb23 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation 2025 , address=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7b62a59c-d62b-40a8-950f-a8e86f33fbef · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation First Conference on Language Modeling , year=
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f402b79-42c9-4c75-bb00-c16d1a5cf483 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM ^2 ) , pages=
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92e8be16-1bd9-4e99-8479-f931ee0db13d · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation NPJ digital medicine , volume=
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7851f010-c71e-4784-819d-5faf2b95d91b · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts , pages=
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0adf9169-21be-4635-b4f3-e681ab698b64 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 2nd Workshop on Human Evaluation of NLP Systems (HumEval) , pages=
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0281dc86-eae2-4c2d-9e48-d6d8d5c9fb87 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the Fourth Workshop on Insights from Negative Results in NLP , pages=
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef034d17-1977-43df-b007-e9d0303743bb · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Computational Linguistics , volume=
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af3084de-54ff-4635-ab0c-ace829bba0a9 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation University of Chicago Coase-Sandor Institute for Law & Economics Research Paper , number=
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdea24ff-4e05-4d6a-8772-a5ece13379bd · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM ^2 ) , pages=
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88298a41-01d8-42aa-8d7b-64f7657b6375 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation LLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4d438264-bbc2-4a75-afce-c3cb7c8db1cf · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Advances in Neural Information Processing Systems , volume=
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7354b67-6b30-476d-8917-49ef0f9f56ea · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation npj Health Systems , volume=
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95367f82-0974-4762-8d4d-4035c3b2d267 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04d912d3-043c-4af2-b405-f0741be2e3df · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcf63442-4f57-444a-9a5a-b5e8ec1ef364 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 31st International Conference on Computational Linguistics , pages=
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfe87851-4caa-42ba-b253-e3910c92ec55 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cd910f03-e810-4dc0-8a36-b4b13e7c6770 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Nature human behaviour , volume=
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31fbb181-f169-4e58-bc82-458b9d881756 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation 2022 , url =
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01f1039a-74f2-43d3-a22e-a95ece5e03fa · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b523a3cc-2551-4327-bd6d-f8f1b0376a0f · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bd170a5-8424-4fbd-9a61-b7a86e51a56e · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation PaperQA: Retrieval-Augmented Generative Agent for Scientific Research
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9c9ebfc8-6d43-4280-bf46-7259751c6732 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation arXiv e-prints , pages=
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca1cb1e1-32dc-436e-97cc-8d24ac2cce54 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4ade0dea-bbf2-4462-b04d-0f42eedda8df · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Text summarization branches out , pages=
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5a05ed5-8ba4-4a53-997e-4f5d9c57bedf · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Mistral 7B
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 17073da2-dec6-4adf-ac80-1ebfb87ec695 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation GPT-4 Technical Report
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3c2823db-2c19-48de-a9f6-3a5156808422 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Clinical Natural Language Processing Workshop , year=
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7d9876d-33c0-44ff-b2e2-f083aac8379b · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Conference on Empirical Methods in Natural Language Processing , year=
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c7ddf9c-ab5e-44ac-b4e5-623e0f1957c7 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Annual Meeting of the Association for Computational Linguistics , year=
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a40d2b5-b97e-49b3-900d-768dd06ffe68 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Annual Meeting of the Association for Computational Linguistics , year=
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d86d1e9b-946b-4b4c-9040-e7ceefc304af · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Annual Meeting of the Association for Computational Linguistics , year=
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfb158c3-27fc-47ef-80ac-c778966136db · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 989d5e63-135a-41d1-8d38-37dea9f1b515 · outbound
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 680c6a2d-57a6-4bd9-8e20-37e1514d6030 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01d7754b-ca66-4846-ab93-90daf0fd7095 · outbound
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d82d8d15-5815-49dc-9a3e-dcef00ea0bd7 · outbound
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e4abbe-3665-4ae9-b218-ec7115280f91 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Yu, Qiang Yang, and Xing Xie
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 79981ea4-a605-417e-bc27-999a4909d88e · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Improving Factuality in Clinical Abstractive Multi-Document Summarization by Guided Continued Pre-training
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3d4ed799-7988-4fa4-af97-67c11843529c · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 50e44a69-99ad-4685-886a-608f25a6e51c · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 42a97c9f-9026-4ebe-aa8f-a0d5c1ea47b9 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8578f1d1-db02-472f-b8a9-4aa23a7a097a · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation medRxiv , pages=
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4fa0a0c-3c0f-48ee-9634-649bcf972adc · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation and Villarroel, Mauricio and Clifford, Gari D
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a75cdf03-cbb2-47ce-a62c-5c715c5b1507 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6910d7ec-6ce3-40ad-9f75-7ac69f68de76 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Bioinformatics , volume =
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e77617e3-a584-442e-8909-e1c024736c16 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation TOPICAL : TOPIC Pages A utomagica L ly
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 500462cd-5edb-4a08-a252-2a017dfee225 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation What’s in a Summary? Laying the Groundwork for Advances in Hospital-Course Summarization
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d629c238-c82e-42a1-83b7-272d3e0ee3b7 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation 2015 26th international workshop on database and expert systems applications (dexa) , pages=
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 096afb44-0e96-4edd-b81d-1725cbfabe3c · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Journal of Intelligent Connectivity and Emerging Technologies , volume=
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c57ae8e-dbbc-4b85-bf9e-d929a1b47e12 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Journal of biomedical informatics , volume=
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7222757e-d931-478a-b6d4-35362ae28b33 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation A Novel System for Extractive Clinical Note Summarization using EHR Data
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 93100199-bc46-4fef-aed2-dcc1c5576915 · outbound
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 02a5fc08-22d1-44df-bc55-00b3c8ab39dc · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Towards Automating Medical Scribing : Clinic Visit D ialogue2 N ote Sentence Alignment and Snippet Summarization
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 35044527-a97a-4540-b321-b02590d9c47a · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation DERA : Enhancing Large Language Model Completions with Dialog-Enabled Resolving Agents
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 66a7e94b-90ac-4811-b637-e256c3f00507 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation JMIR medical education , volume=
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3858b38e-2fb3-4e28-b94a-03ee0238e5df · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Cureus , volume=
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1d50d2e-fa1b-4217-9241-dbd2f66e2458 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83ced66a-83d2-456d-86b9-966a6712b6c2 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation From Protocol to Screening: A Hybrid Learning Approach for Technology-Assisted Systematic Literature Reviews
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 499f5a0b-782a-4ff5-bc47-dfb225598d0a · outbound
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3d50d6e7-656c-4b2d-b1d7-6fe1267d4be4 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation On Learning to Summarize with Large Language Models as References
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d009cb8c-f181-432d-a6d7-9cfca7c6bfc0 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Summarizing, simplifying, and synthesizing medical evidence using GPT -3 (with varying success)
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d4d46906-9517-447a-802f-13ff88d86416 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation An empirical survey on long document summarization: Datasets, models, and metrics
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f4033730-755b-4562-918d-2a7897fc2b42 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Conference on Empirical Methods in Natural Language Processing , year=
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf87f024-4abc-49d1-928d-8dc81c87a1e7 · outbound
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c59feabe-6e95-4cd8-955d-553e116ae950 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation 2023 , journal =
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3525dbf9-67e6-4d0f-9ca2-c17c250cca5c · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Annual Meeting of the Association for Computational Linguistics , year=
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 058447c0-3117-4e47-bf33-1e40dbf769ef · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df9f31c7-484f-4912-969d-a97d9286169e · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Journal of Medical Internet Research , year=
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23a06557-3ed7-47ad-8984-5345d7388e81 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation BMJ Open , year=
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e273582a-ea3e-4508-98b3-f36c2d1b034b · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Energies , year=
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71b6e608-f57d-4808-904c-a7a0a6935368 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 52nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29700864-7123-40c8-a28a-0287941ffb0e · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Unresolved cited work
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fdf8532-0405-4246-97b5-f7d5d0c3c282 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the conference
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0844e4b-fb9c-4de3-a675-905732047568 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Trends in cognitive sciences , volume=
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a318ff8-fdff-459b-aea2-26b4ea2b1a21 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Craik , abstract =
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bdf59089-fc1d-4dc1-861a-6b7b06fcdbe2 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Conference on Empirical Methods in Natural Language Processing , year=
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation add8bb77-99b8-44f8-a243-419ffef34766 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8041b7b-47db-417a-8afc-c4eb9c13c160 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40afc49a-429e-48b5-a658-cfab1b4c2c3e · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Annual Meeting of the Association for Computational Linguistics , year=
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19d1e942-015b-4bf7-8162-42ab26d79b24 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , year=
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec1320b5-868c-48f4-b46d-4da9355d15af · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation ArXiv , year=
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d914574-9092-439e-9e21-c88e7c18bdad · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation H i S truct+: Improving Extractive Text Summarization with Hierarchical Structure Information
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 997af391-06bd-48f0-a7e6-ce3fda19e9c1 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation What factors might be causing the significant deviations in my circadian rhythm patterns over the past 30 days?
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2418d2f6-894b-4402-835b-8e4857e8bdbc · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation and Belz, Anya and Du s ek, Ond r ej and Mille, Simon and van der Lee, Chris and Reiter, Ehud and Santhanam, Shivani and Thomson, Craig
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation addd4d8d-530d-45cd-8296-1f1f828e1ab5 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation G-Eval : NLG Evaluation using GPT-4 with Better Human Alignment
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81dc610d-89e7-45e1-b7e4-cf622802c639 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Governance Framework
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b8eb6e69-c084-4c87-b914-610b0f4c24ba · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2cabe6f7-8e1a-4cda-abb1-132bd67c0256 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Findings of the Association for Computational Linguistics: EMNLP 2023 , year=
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eac47b5-1fd2-427e-9246-ef72f0dda100 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d52780b4-b71b-4926-aea5-9ddee4067efc · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation be54c590-61bb-49ff-b72c-860402bd6037 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Transactions of the Association for Computational Linguistics , volume=
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca0ce1c-c752-4e57-a8d3-76d7e35c5565 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation FFCI: A Framework for Interpretable Automatic Evaluation of Summarization
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e2964047-2545-45f2-aa8e-40f212b2f934 · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Journal of the Royal Society of Medicine , volume =
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 556ded4f-87cc-4e5c-97f3-cf358ea1493c · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation The significant cost of systematic reviews and meta-analyses: A call for greater involvement of machine learning to assess the promise of clinical trials , journal =
Reference 99
Source-reported events for the cited work
correction dated 2019-09-12. Source: crossref record 10.1016/j.conctc.2019.100450->10.1016/j.conctc.2019.100443:correction, observed 2026-07-11T03:16:42.407043+00:00. This notice travels one citation hop only.
Observation 7a04601b-ddd8-4b7d-9cd9-903a4cb5e19e · outbound
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation npj Digital Medicine , volume=
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.