Pith. sign in

REVIEW 9 cited by

HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.07833 v1 pith:CVONFEY4 submitted 2025-03-10 cs.CL cs.AI

HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations

classification cs.CL cs.AI
keywords hallucinationsdatasetfine-grainedhalluverse25multilingualcategorizescontextshallucination
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large Language Models (LLMs) are increasingly used in various contexts, yet remain prone to generating non-factual content, commonly referred to as "hallucinations". The literature categorizes hallucinations into several types, including entity-level, relation-level, and sentence-level hallucinations. However, existing hallucination datasets often fail to capture fine-grained hallucinations in multilingual settings. In this work, we introduce HalluVerse25, a multilingual LLM hallucination dataset that categorizes fine-grained hallucinations in English, Arabic, and Turkish. Our dataset construction pipeline uses an LLM to inject hallucinations into factual biographical sentences, followed by a rigorous human annotation process to ensure data quality. We evaluate several LLMs on HalluVerse25, providing valuable insights into how proprietary models perform in detecting LLM-generated hallucinations across different contexts.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HalluScore: Large Language Model Hallucination Question Answering Benchmark

    cs.CL 2026-05 unverdicted novelty 7.0

    HalluScore is a curated Arabic QA dataset with 827 questions, ground-truth evidence, and human annotations used to measure hallucination rates across 17 LLMs.

  2. Harnessing Embodied Agents: Runtime Governance for Policy-Constrained Execution

    cs.RO 2026-04 unverdicted novelty 7.0

    A runtime governance framework for embodied agents achieves 96.2% interception of unauthorized actions and 91.4% recovery success in 1000 simulation trials by externalizing policy enforcement.

  3. HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

    cs.CL 2026-07 conditional novelty 6.0

    HalluTruthQA provides 2,400 expert-annotated Arabic QA examples with character-level hallucination spans, explanations, and verification candidates, and shows no single LLM excels at all four evaluation tasks.

  4. HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

    cs.CL 2026-07 conditional novelty 6.0

    HalluTruthQA contributes 2,400 expert-annotated Arabic QA examples with hallucination labels, error spans, human explanations, and candidate answers; evaluations show no open LLM leads across all four tasks.

  5. Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering

    cs.CL 2026-07 conditional novelty 6.0

    Prompt-point activations carry a graded, steerable entity-familiarity signal that is robust to Polish/English stem changes and is stronger in Polish-adapted models than in base models.

  6. Does Bielik Know What It Doesn't Know? Activation Dispersion Separates Entity Familiarity from Factual Reliability Across Model Scale

    cs.CL 2026-07 conditional novelty 6.0

    Unsupervised MLP activation dispersion separates known from fabricated entities at AUROC 0.95–1.00 across Bielik scales, while factual reliability scales separately and refusals stay near zero.

  7. PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations

    cs.CL 2026-04 unverdicted novelty 6.0

    PRISM benchmark disentangles LLM hallucinations into knowledge missing, knowledge errors, reasoning errors, and instruction-following errors across three generation stages, revealing trade-offs when testing 24 models.

  8. Harnessing Embodied Agents: Runtime Governance for Policy-Constrained Execution

    cs.RO 2026-04 conditional novelty 5.5

    An external runtime governance layer for embodied agents intercepts unauthorized actions at ~96% and recovers from runtime drift at ~91% under policy constraints in simulation, outperforming pre-execution-only baselin...

  9. Harnessing Embodied Agents: Runtime Governance for Policy-Constrained Execution

    cs.RO 2026-04 unverdicted novelty 5.0

    A runtime governance framework for embodied agents intercepts 96.2% of unauthorized actions and achieves 91.4% recovery success in 1000 simulation trials while outperforming baselines.