Pith. sign in

Paper Citation Record · LEDGER

An Empirical Study of Evaluating Long-form Question Answering

As of 16 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2504.18413.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18413 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:21:26.126300Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved48
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 58b8a280-9297-446e-a097-2d386f68af59 · outbound

This paper cites Can we trust the evaluation on ChatGPT?.

An Empirical Study of Evaluating Long-form Question Answering Can we trust the evaluation on ChatGPT?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.883260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.883260Z digest=sha256:c6c73d6f1baec38386b1c1e47405bd1decdab6ee7ffc93f2cb44f1321974bb94

Observation 97a8bd43-5ede-4d9c-b116-a90c1d2ce30f · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:21:27.125266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T10:21:25.888983Z digest=sha256:29cc9f9d55ca78c6abcdaace5f06522742ffa99468d1955e4b024aa114fb245e

Observation f0ae7b8f-427a-48e3-ba4c-eea5b9fb9c1f · outbound

This paper cites Bruce Croft, and Mark Sanderson.

An Empirical Study of Evaluating Long-form Question Answering Bruce Croft, and Mark Sanderson

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.898236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.898236Z digest=sha256:6a486a0fedb21313b2e7d5ff81724107d63a063541788664676d32c0a6a96ded

Observation 63da9470-c3e5-4314-bb46-260aa2133211 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.903209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.903209Z digest=sha256:f9ca2839af2c0463b6c6b14ad16ff2c31295e16c31016ba4159647a4d80d6e0a

Observation 970a9468-29f8-4830-818e-cfc359ebefe7 · outbound

This paper cites Exploring the Use of Large Language Models for Reference-Free Text Quality Evaluation: An Empirical Study.

An Empirical Study of Evaluating Long-form Question Answering Exploring the Use of Large Language Models for Reference-Free Text Quality Evaluation: An Empirical Study

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.911997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.911997Z digest=sha256:ee1cdaa56400ce845c91edddb66d38b6c7e6af3f1c0726a0c696fd1fbcb6f6a8

Observation d9cf9ec3-2a24-4fbd-b013-697198005c01 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

An Empirical Study of Evaluating Long-form Question Answering Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.907544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.907544Z digest=sha256:aac552d9ce4ba4a179c813b95c1a69480f3d62553d8e22f62dfb817fbd742932

Observation 47647def-fcef-4404-9f48-38fca9d1d730 · outbound

This paper cites A Closer Look into Automatic Evaluation Using Large Language Models.

An Empirical Study of Evaluating Long-form Question Answering A Closer Look into Automatic Evaluation Using Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.922444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.922444Z digest=sha256:a074186c6cb95ca6cc8ca229ce376ac14e24873f4b9774ba0ac195c0d40daeec

Observation e8da3c36-ba94-4cdb-a8e2-f25f18e0f5db · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

An Empirical Study of Evaluating Long-form Question Answering Can Large Language Models Be an Alternative to Human Evaluations?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.917092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.917092Z digest=sha256:eb1396a5a03f9f33fc2d8360de879ff0e57f37bca6074cbe9453681bc76960b8

Observation b3a24183-b2d9-4a6a-830a-25da1f545b31 · outbound

This paper cites Ragas: Automated Evaluation of Retrieval Augmented Generation.

An Empirical Study of Evaluating Long-form Question Answering Ragas: Automated Evaluation of Retrieval Augmented Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.932874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.932874Z digest=sha256:a202c51db0115b66e876d37b61831b831e217a44fcfc6599d652a2841905cd3b

Observation f21e1034-5f70-42bb-b0e9-c9226d7424db · outbound

This paper cites On The Evaluation of Machine Translation Systems Trained With Back-Translation.

An Empirical Study of Evaluating Long-form Question Answering On The Evaluation of Machine Translation Systems Trained With Back-Translation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:21:26.309260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T10:21:25.927966Z digest=sha256:1150c1c36546cdd18ced086de1cea3fbfb20ab3db56c3adc7c122106c154b146

Observation ceb0ed0c-80db-4da4-8097-dad59deb5be9 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-16T10:21:26.790440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T10:21:25.941380Z digest=sha256:eda9ac01e0e3bdaa789a40cbbb667820044ebdaa957e64605fa6f4d77d3a0714

Observation ee6ea400-a4ec-4edd-9333-3864238fc704 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:21:27.100410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T10:21:25.937193Z digest=sha256:24ea7771834ec2b32aad0fdb4a4b4a982aa54b7c035e70de6bda2a01ffdddf75

Observation 6460a8a6-db23-4f55-9c46-50894b02ccc3 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.949688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.949688Z digest=sha256:94535a669e44e3e1bad4f6b7d7a876622c85dfe7c405fa5347d3f26941d39d86

Observation e3300c13-4dac-47fb-98d9-aa212b83fe8f · outbound

This paper cites GPTScore: Evaluate as You Desire.

An Empirical Study of Evaluating Long-form Question Answering GPTScore: Evaluate as You Desire

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.945443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.945443Z digest=sha256:e131408786fc9615b84f253aa58d98c16e00a5988861be21114adc0d11d7e499

Observation efde0744-6a09-40d8-bd32-26341c908e18 · outbound

This paper cites A Survey on LLM-as-a-Judge.

An Empirical Study of Evaluating Long-form Question Answering A Survey on LLM-as-a-Judge

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.963232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.963232Z digest=sha256:e7fac0788b770cd8c0da04cf905ce2fea5451d0ce0df701b2e881eccb4b767b5

Observation 74114987-1bfc-4a8f-8686-0661a60d711f · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:21:27.075614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T10:21:25.968096Z digest=sha256:26e35204b9d0096562eebd698e9b0f5e51a26731fd5aa1736876304700211d6c

Observation 29a718aa-c2b7-4a10-9410-1119160c8518 · outbound

This paper cites Translationese in Machine Translation Evaluation.

An Empirical Study of Evaluating Long-form Question Answering Translationese in Machine Translation Evaluation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.958494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.958494Z digest=sha256:720d6252a9c496e995232409dbc2cd0793c05c2b8d1b6874e2cb971094d7088c

Observation 239da420-edc5-4774-92be-a70179a5704e · outbound

This paper cites Mistral 7B.

An Empirical Study of Evaluating Long-form Question Answering Mistral 7B

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.981933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.981933Z digest=sha256:bf2c5c78b4eb63d3d7e9ebd6caab117df83359d3d2ee74bf0f3fe86fea8e4f3c

Observation 82c9c9e2-c124-4501-9848-8f46cf12dc8e · outbound

This paper cites SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-Scaling.

An Empirical Study of Evaluating Long-form Question Answering SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-Scaling

Reference 19

Resolution
malformed identifier
no resolver link, observed 2026-08-16T10:21:25.986805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.986805Z digest=sha256:4cf01d1e31e98da3560be4cac3297615e63c2419883a0e3a33ee7d3d58f77fc6

Observation 0efe1c4b-d7b0-4ddd-a419-36f3eb51cfaf · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.991809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.991809Z digest=sha256:81721910f72602d5895b4682d344719d8138b7e0e155c25930cbec15c0b88cf5

Observation 4d31df53-2e51-423e-8943-b07bd5b4e0c3 · outbound

This paper cites Survey of Hallucination in Natural Language Generation.

An Empirical Study of Evaluating Long-form Question Answering Survey of Hallucination in Natural Language Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.977125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.977125Z digest=sha256:0e998a09bb7db9353f70042793846b3d5446fd2e873e8b84e5e6dc2bb7e25cf4

Observation 870190f9-74a1-45c8-8fbe-e05ab7516e85 · outbound

This paper cites Manning, Christopher Ré, Diana Acosta-Navas, Drew A.

An Empirical Study of Evaluating Long-form Question Answering Manning, Christopher Ré, Diana Acosta-Navas, Drew A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:21:27.036109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T10:21:26.000929Z digest=sha256:347d026a144ce88e64e922e11cb1a858b3afaf6537cae94958f93c3e8a8e858d

Observation 93d24641-1b2c-4c0d-aaca-6858d73aca31 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.010113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.010113Z digest=sha256:616337fd740a8a0b2151c41799979b4c880abd678c0c68024d7e9cf2394ff070

Observation 0959be33-1ff3-4d99-bf62-e0700bccf69b · outbound

This paper cites LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models.

An Empirical Study of Evaluating Long-form Question Answering LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.014390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.014390Z digest=sha256:ccb3703d25ec2d89dea40fdb4745611828f2d88eb038ac39fbff1c27725bed5d

Observation ee5747c1-46fa-4e75-8990-c261c810452c · outbound

This paper cites A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets.

An Empirical Study of Evaluating Long-form Question Answering A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.996298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.996298Z digest=sha256:2404fec4a7e62c5e1d16f47c3164b5f1bac312fef7aee91dd10e6742a93ecd20

Observation 35823cd7-a060-4ab9-aea8-d1ab449fe8db · outbound

This paper cites LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models.

An Empirical Study of Evaluating Long-form Question Answering LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.022925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.022925Z digest=sha256:733c034871e24203c6133ad38bba524161da8c2593913a557ef0e9e9f7aea3da

Observation 9b225c7c-cd45-4633-8dac-12ecded4109d · outbound

This paper cites Language Models are Few-Shot Learners.

An Empirical Study of Evaluating Long-form Question Answering Language Models are Few-Shot Learners

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.027310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.027310Z digest=sha256:270d6039b61c4ce3ecaafab61fd976cd7e97e74f44a99038caa512556134fdf4

Observation bb4d1a79-bfbd-4e6a-b51d-b83065467d81 · outbound

This paper cites FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation.

An Empirical Study of Evaluating Long-form Question Answering FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.031803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.031803Z digest=sha256:e942cc36f1afa5c4656c2d7c371a298327f48c04a732b1d269c52941c4bf8ec2

Observation 4d529b29-d329-478b-a4fc-f0d04736caa1 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:21:27.011563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T10:21:26.036149Z digest=sha256:e74b4503b8518abc47c7f8b8c33d9db12830d7d21212e45e5001d52aad8a9b24

Observation a471f6db-7abd-4408-841f-be27588aef51 · outbound

This paper cites G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment.

An Empirical Study of Evaluating Long-form Question Answering G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.018557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.018557Z digest=sha256:4ce6d5de18f9acbf50fa81689da8f89624341a094626f8f84a46d2df1c92de7e

Observation 5632106d-e38c-4b74-adf5-33ced804eee7 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.044346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.044346Z digest=sha256:8c2c09ac710e51d46c6f4e08b6bbf22833841f615ec28699907de72afdf8e12a

Observation 4956660f-98d4-4e57-9a76-0b9068815c59 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 32

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T10:21:26.971017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T10:21:26.048915Z digest=sha256:82e4d8a614b0ac7b46aace193de7de148b2f216e6180cb03c94bfc71f746d6a4

Observation 19e2b9e0-5e67-48e0-9f8d-d34f9cdc8cd9 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.053539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.053539Z digest=sha256:efb807cd7af26bfb01f42597a926a761bd9b681e99ed204e401799f6eed4d405

Observation df03481d-2457-44a0-a66c-478b518706a2 · outbound

This paper cites Read before Generate! Faithful Long Form Question Answering with Machine Reading.

An Empirical Study of Evaluating Long-form Question Answering Read before Generate! Faithful Long Form Question Answering with Machine Reading

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.058961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.058961Z digest=sha256:39df7b7c3a1a381abcb640f2c58889ed44ca64f26b2f5ba740a1fcc563371183

Observation f5619ef3-05b5-4d24-a1a7-6164bcab6724 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:21:26.996376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T10:21:26.040419Z digest=sha256:541d8316924bb546a800c9b6739a1a65ae0f6046875efa56fe0e8408042c5a4e

Observation 464041cb-7668-40bc-ab32-c4c800eb77a6 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:21:26.955258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T10:21:26.067967Z digest=sha256:689f30c30c6937618ad8a986b66d024078213a49fc97b1d8984b64018238bc23

Observation 7b2978f5-28f0-426e-aaa8-e1a7e330ecdc · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

An Empirical Study of Evaluating Long-form Question Answering Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.072226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.072226Z digest=sha256:f88c79a8cda639160fb4fbcf7e40af5b807d2de4fa04000fe240a2be0901cf36

Observation d495196b-ab18-4795-94f8-1b9ed58bd2d9 · outbound

This paper cites Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation.

An Empirical Study of Evaluating Long-form Question Answering Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.076185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.076185Z digest=sha256:b458ef37f6a0b6bba1519e08c8f1c9a94442d96f3ad16b4ecb61f0f71ee2bb36

Observation dee73ccb-8926-4517-a994-d3ec9cc72a61 · outbound

This paper cites Is ChatGPT a Good NLG Evaluator? A Preliminary Study.

An Empirical Study of Evaluating Long-form Question Answering Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.080417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.080417Z digest=sha256:1aabb131d175ed47cf67b77d5d56eed85af266b29a195fc8575e44a06a5e4c0a

Observation c9dd1258-4c87-445d-8027-3d947f69c39a · outbound

This paper cites Chain-of-Discussion: A Multi-Model Framework for Complex Evidence-Based Question Answering.

An Empirical Study of Evaluating Long-form Question Answering Chain-of-Discussion: A Multi-Model Framework for Complex Evidence-Based Question Answering

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.063537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.063537Z digest=sha256:3ee76f673ee035998d440a35f1a16d34d601530b9df78183adb628425d32c294

Observation ee0e2d8b-5bfd-4410-a804-db642cd30dc2 · outbound

This paper cites PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization.

An Empirical Study of Evaluating Long-form Question Answering PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.088890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.088890Z digest=sha256:ff050b3f2c1cc2fb4eb6332309edb308c698023bdbc212ea8d99845e8a980eaa

Observation c672f35b-7bcd-4386-b51f-3e89714af8f6 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:21:26.939641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T10:21:26.093480Z digest=sha256:505f31cd5ae72972f7f3f2e99ab860aad87e76930af6d0fb861fabff58413e3c

Observation 23252623-4f54-49d5-a877-3f665cc40096 · outbound

This paper cites A Critical Evaluation of Evaluations for Long-form Question Answering.

An Empirical Study of Evaluating Long-form Question Answering A Critical Evaluation of Evaluations for Long-form Question Answering

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.097838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.097838Z digest=sha256:44fd2bf8662fc847cc6714730de431739b253d79c668f85f41550683c98582aa

Observation 56a860e5-b077-4fab-a327-b6b3f30c6384 · outbound

This paper cites Prompt Engineering a Prompt Engineer.

An Empirical Study of Evaluating Long-form Question Answering Prompt Engineering a Prompt Engineer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.102266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.102266Z digest=sha256:de1288136266c05a42881c26427205febab8856d6fc5773189a0048e5846c5a3

Observation 8b488337-628e-4615-87b9-bc9db51940b7 · outbound

This paper cites Large Language Models are not Fair Evaluators.

An Empirical Study of Evaluating Long-form Question Answering Large Language Models are not Fair Evaluators

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.084583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.084583Z digest=sha256:f2215e0cd5d7cd24374a0bd551d7c0d88992af5d169d857ce5a4a95352dc5e5f

Observation 4d74b7f9-b248-40e0-b4dc-276ae4091f6d · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

An Empirical Study of Evaluating Long-form Question Answering BERTScore: Evaluating Text Generation with BERT

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.112703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.112703Z digest=sha256:624bd465f2873ff7f2a53b1a1ab7bb8c59b7cd612f56554aea65c853a9c44d76

Observation 75241382-06bd-4b0c-9af0-5f7d869d76bf · outbound

This paper cites LLMEval: A Preliminary Study on How to Evaluate Large Language Models.

An Empirical Study of Evaluating Long-form Question Answering LLMEval: A Preliminary Study on How to Evaluate Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.117543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.117543Z digest=sha256:f1e593775ba5a5397e259dac89fccb53bfd6c59479e707f870a2ed0f34c371b5

Observation d1fbc60f-3d97-43af-b20c-3fd2ff306ce8 · outbound

This paper cites Meyer, and Stef- fen Eger.

An Empirical Study of Evaluating Long-form Question Answering Meyer, and Stef- fen Eger

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.122295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.122295Z digest=sha256:5223f45946e6e065a1c56c1130cc9f660e8b29788f8846dde7a9408be7a09498

Observation d823a645-2ceb-4416-b70a-a68d3e580128 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Evaluating Long-form Question Answering Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:21:26.913420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T10:21:26.126300Z digest=sha256:764cf7c7dad453f402fe7b09b08ff8b4563bbf689c1e18b17dc176463e167310

Observation 84a6d43c-35b1-476d-aa3a-e21e8df48cbc · outbound

This paper cites FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets.

An Empirical Study of Evaluating Long-form Question Answering FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.107650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.107650Z digest=sha256:ca098fed559167e8be06930982235104195b48781858e1d4fc00bd4e68d514f6

Observation fd397251-7d15-4dde-94fb-5ad550defb35 · outbound

This paper cites In Advances in Information Retrieval: 42nd European Conference on IR Research, ECIR 2020, Lisbon, Portugal, April 14–17, 2020, Proceedings, Part II 42.

An Empirical Study of Evaluating Long-form Question Answering In Advances in Information Retrieval: 42nd European Conference on IR Research, ECIR 2020, Lisbon, Portugal, April 14–17, 2020, Proceedings, Part II 42

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:21:27.061012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T10:21:25.972464Z digest=sha256:be202dd7ca539513ec7ef1dc7bbb98a5e243619df1b477bd4517fb9d5175d6a6

Observation 4e61d8c8-5d00-4acf-be30-a673f89df2b3 · outbound

This paper cites Holistic Evaluation of Language Models.

An Empirical Study of Evaluating Long-form Question Answering Holistic Evaluation of Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.005837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.005837Z digest=sha256:18dfafaad35d89b61cd5d380462d032e44feb30b0b7abade3539255d260b3ba9

Observation b468ece7-0fa4-4138-942d-721e0eb5f7a2 · outbound

This paper cites Investigating Answerability of LLMs for Long-Form Question Answering.

An Empirical Study of Evaluating Long-form Question Answering Investigating Answerability of LLMs for Long-Form Question Answering

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.893306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.893306Z digest=sha256:d4c9c67e2b8d5c2b0050eb162df93afd6f424622011b31b603141273ec27e3be

Observation c93eb2fc-df8a-430a-843e-2975a7da8c10 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

An Empirical Study of Evaluating Long-form Question Answering ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.954165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.954165Z digest=sha256:456b44b630d76023a02330f93d6a1152948a8cbeb971ab47cb2635a2a7ba1a80

Pith citing papers

No inbound Pith citation observations are available.