Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:21:26.126300Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2504.18413.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:21:26.126300Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 58b8a280-9297-446e-a097-2d386f68af59 · outbound
An Empirical Study of Evaluating Long-form Question Answering Can we trust the evaluation on ChatGPT?
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97a8bd43-5ede-4d9c-b116-a90c1d2ce30f · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f0ae7b8f-427a-48e3-ba4c-eea5b9fb9c1f · outbound
An Empirical Study of Evaluating Long-form Question Answering Bruce Croft, and Mark Sanderson
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63da9470-c3e5-4314-bb46-260aa2133211 · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 970a9468-29f8-4830-818e-cfc359ebefe7 · outbound
An Empirical Study of Evaluating Long-form Question Answering Exploring the Use of Large Language Models for Reference-Free Text Quality Evaluation: An Empirical Study
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9cf9ec3-2a24-4fbd-b013-697198005c01 · outbound
An Empirical Study of Evaluating Long-form Question Answering Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47647def-fcef-4404-9f48-38fca9d1d730 · outbound
An Empirical Study of Evaluating Long-form Question Answering A Closer Look into Automatic Evaluation Using Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8da3c36-ba94-4cdb-a8e2-f25f18e0f5db · outbound
An Empirical Study of Evaluating Long-form Question Answering Can Large Language Models Be an Alternative to Human Evaluations?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3a24183-b2d9-4a6a-830a-25da1f545b31 · outbound
An Empirical Study of Evaluating Long-form Question Answering Ragas: Automated Evaluation of Retrieval Augmented Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f21e1034-5f70-42bb-b0e9-c9226d7424db · outbound
An Empirical Study of Evaluating Long-form Question Answering On The Evaluation of Machine Translation Systems Trained With Back-Translation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ceb0ed0c-80db-4da4-8097-dad59deb5be9 · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ee6ea400-a4ec-4edd-9333-3864238fc704 · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6460a8a6-db23-4f55-9c46-50894b02ccc3 · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3300c13-4dac-47fb-98d9-aa212b83fe8f · outbound
An Empirical Study of Evaluating Long-form Question Answering GPTScore: Evaluate as You Desire
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efde0744-6a09-40d8-bd32-26341c908e18 · outbound
An Empirical Study of Evaluating Long-form Question Answering A Survey on LLM-as-a-Judge
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74114987-1bfc-4a8f-8686-0661a60d711f · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 29a718aa-c2b7-4a10-9410-1119160c8518 · outbound
An Empirical Study of Evaluating Long-form Question Answering Translationese in Machine Translation Evaluation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 239da420-edc5-4774-92be-a70179a5704e · outbound
An Empirical Study of Evaluating Long-form Question Answering Mistral 7B
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82c9c9e2-c124-4501-9848-8f46cf12dc8e · outbound
An Empirical Study of Evaluating Long-form Question Answering SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-Scaling
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0efe1c4b-d7b0-4ddd-a419-36f3eb51cfaf · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d31df53-2e51-423e-8943-b07bd5b4e0c3 · outbound
An Empirical Study of Evaluating Long-form Question Answering Survey of Hallucination in Natural Language Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 870190f9-74a1-45c8-8fbe-e05ab7516e85 · outbound
An Empirical Study of Evaluating Long-form Question Answering Manning, Christopher Ré, Diana Acosta-Navas, Drew A
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 93d24641-1b2c-4c0d-aaca-6858d73aca31 · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0959be33-1ff3-4d99-bf62-e0700bccf69b · outbound
An Empirical Study of Evaluating Long-form Question Answering LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee5747c1-46fa-4e75-8990-c261c810452c · outbound
An Empirical Study of Evaluating Long-form Question Answering A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35823cd7-a060-4ab9-aea8-d1ab449fe8db · outbound
An Empirical Study of Evaluating Long-form Question Answering LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b225c7c-cd45-4633-8dac-12ecded4109d · outbound
An Empirical Study of Evaluating Long-form Question Answering Language Models are Few-Shot Learners
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb4d1a79-bfbd-4e6a-b51d-b83065467d81 · outbound
An Empirical Study of Evaluating Long-form Question Answering FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d529b29-d329-478b-a4fc-f0d04736caa1 · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a471f6db-7abd-4408-841f-be27588aef51 · outbound
An Empirical Study of Evaluating Long-form Question Answering G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5632106d-e38c-4b74-adf5-33ced804eee7 · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4956660f-98d4-4e57-9a76-0b9068815c59 · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 19e2b9e0-5e67-48e0-9f8d-d34f9cdc8cd9 · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df03481d-2457-44a0-a66c-478b518706a2 · outbound
An Empirical Study of Evaluating Long-form Question Answering Read before Generate! Faithful Long Form Question Answering with Machine Reading
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5619ef3-05b5-4d24-a1a7-6164bcab6724 · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 464041cb-7668-40bc-ab32-c4c800eb77a6 · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7b2978f5-28f0-426e-aaa8-e1a7e330ecdc · outbound
An Empirical Study of Evaluating Long-form Question Answering Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d495196b-ab18-4795-94f8-1b9ed58bd2d9 · outbound
An Empirical Study of Evaluating Long-form Question Answering Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dee73ccb-8926-4517-a994-d3ec9cc72a61 · outbound
An Empirical Study of Evaluating Long-form Question Answering Is ChatGPT a Good NLG Evaluator? A Preliminary Study
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9dd1258-4c87-445d-8027-3d947f69c39a · outbound
An Empirical Study of Evaluating Long-form Question Answering Chain-of-Discussion: A Multi-Model Framework for Complex Evidence-Based Question Answering
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee0e2d8b-5bfd-4410-a804-db642cd30dc2 · outbound
An Empirical Study of Evaluating Long-form Question Answering PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c672f35b-7bcd-4386-b51f-3e89714af8f6 · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 23252623-4f54-49d5-a877-3f665cc40096 · outbound
An Empirical Study of Evaluating Long-form Question Answering A Critical Evaluation of Evaluations for Long-form Question Answering
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56a860e5-b077-4fab-a327-b6b3f30c6384 · outbound
An Empirical Study of Evaluating Long-form Question Answering Prompt Engineering a Prompt Engineer
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b488337-628e-4615-87b9-bc9db51940b7 · outbound
An Empirical Study of Evaluating Long-form Question Answering Large Language Models are not Fair Evaluators
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d74b7f9-b248-40e0-b4dc-276ae4091f6d · outbound
An Empirical Study of Evaluating Long-form Question Answering BERTScore: Evaluating Text Generation with BERT
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75241382-06bd-4b0c-9af0-5f7d869d76bf · outbound
An Empirical Study of Evaluating Long-form Question Answering LLMEval: A Preliminary Study on How to Evaluate Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1fbc60f-3d97-43af-b20c-3fd2ff306ce8 · outbound
An Empirical Study of Evaluating Long-form Question Answering Meyer, and Stef- fen Eger
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d823a645-2ceb-4416-b70a-a68d3e580128 · outbound
An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 84a6d43c-35b1-476d-aa3a-e21e8df48cbc · outbound
An Empirical Study of Evaluating Long-form Question Answering FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd397251-7d15-4dde-94fb-5ad550defb35 · outbound
An Empirical Study of Evaluating Long-form Question Answering In Advances in Information Retrieval: 42nd European Conference on IR Research, ECIR 2020, Lisbon, Portugal, April 14–17, 2020, Proceedings, Part II 42
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4e61d8c8-5d00-4acf-be30-a673f89df2b3 · outbound
An Empirical Study of Evaluating Long-form Question Answering Holistic Evaluation of Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b468ece7-0fa4-4138-942d-721e0eb5f7a2 · outbound
An Empirical Study of Evaluating Long-form Question Answering Investigating Answerability of LLMs for Long-Form Question Answering
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c93eb2fc-df8a-430a-843e-2975a7da8c10 · outbound
An Empirical Study of Evaluating Long-form Question Answering ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.