Pith. sign in

REVIEW 1 cited by

Multi-FAct: Assessing Factuality of Multilingual LLMs using FActScore

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.18045 v3 pith:W37IQ7BE submitted 2024-02-28 cs.CL

classification cs.CL
keywords multilingualfactualityevaluationfactualfactscoregenerationllmslong-form
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Evaluating the factuality of long-form large language model (LLM)-generated text is an important challenge. Recently there has been a surge of interest in factuality evaluation for English, but little is known about the factuality evaluation of multilingual LLMs, specially when it comes to long-form generation. %This paper systematically evaluates multilingual LLMs' factual accuracy across languages and geographic regions. We introduce a simple pipeline for multilingual factuality evaluation, by applying FActScore (Min et al., 2023) for diverse languages. In addition to evaluating multilingual factual generation, we evaluate the factual accuracy of long-form text generation in topics that reflect regional diversity. We also examine the feasibility of running the FActScore pipeline using non-English Wikipedia and provide comprehensive guidelines on multilingual factual evaluation for regionally diverse topics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A new seven-language ophthalmology benchmark shows LLMs are less accurate in LMIC languages, and an agentic translation-plus-RAG pipeline reduces the gap.

Pith tools