Pith. sign in

REVIEW 3 cited by

Comparing Hallucination Detection Metrics for Multilingual Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.10496 v2 pith:YVDIJ5SY submitted 2024-02-16 cs.CL cs.AI

classification cs.CLcs.AI
keywords metricsdetectionhallucinationwellhallucinationslanguagesmultilingualhuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While many hallucination detection techniques have been evaluated on English text, their effectiveness in multilingual contexts remains unknown. This paper assesses how well various factual hallucination detection metrics (lexical metrics like ROUGE and Named Entity Overlap, and Natural Language Inference (NLI)-based metrics) identify hallucinations in generated biographical summaries across languages. We compare how well automatic metrics correlate to each other and whether they agree with human judgments of factuality. Our analysis reveals that while the lexical metrics are ineffective, NLI-based metrics perform well, correlating with human annotations in many settings and often outperforming supervised models. However, NLI metrics are still limited, as they do not detect single-fact hallucinations well and fail for lower-resource languages. Therefore, our findings highlight the gaps in exisiting hallucination detection methods for non-English languages and motivate future research to develop more robust multilingual detection methods for LLM hallucinations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention

    cs.CL 2025-05 conditional novelty 6.0 of 10

    CausalAbstain filters multilingual self-feedback by comparing how much it changes the model's abstention decision, improving abstention accuracy over baselines on two benchmarks.

  2. Beyond English: Unveiling Multilingual Bias in LLM Copyright Compliance

    cs.CY 2025-02 conditional novelty 6.0 of 10

    LLMs refuse to reproduce copyrighted song lyrics and reveal verbatim lyrics at very different rates depending on the language of the song and the language of the prompt.

  3. MSA at SemEval-2025 Task 3: High Quality Weak Labeling and LLM Ensemble Verification for Multilingual Hallucination Detection

    cs.CL 2025-05 reject novelty 4.0 of 10

    An LLM ensemble that extracts hallucinated spans and votes on them, followed by fuzzy matching, achieved top ranks in Arabic and Basque at the Mu-SHROOM shared task.

Pith tools