Pith. sign in

REVIEW 3 cited by

SINdex: Semantic INconsistency Index for Hallucination Detection in LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.05980 v1 pith:CGDCJ3U3 submitted 2025-03-07 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsdetectionhallucinationacrossclusteringframeworkinconsistencysemantic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) are increasingly deployed across diverse domains, yet they are prone to generating factually incorrect outputs - commonly known as "hallucinations." Among existing mitigation strategies, uncertainty-based methods are particularly attractive due to their ease of implementation, independence from external data, and compatibility with standard LLMs. In this work, we introduce a novel and scalable uncertainty-based semantic clustering framework for automated hallucination detection. Our approach leverages sentence embeddings and hierarchical clustering alongside a newly proposed inconsistency measure, SINdex, to yield more homogeneous clusters and more accurate detection of hallucination phenomena across various LLMs. Evaluations on prominent open- and closed-book QA datasets demonstrate that our method achieves AUROC improvements of up to 9.3% over state-of-the-art techniques. Extensive ablation studies further validate the effectiveness of each component in our framework.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Theorem-of-Thought: A Multi-Agent Framework for Abductive, Deductive, and Inductive Reasoning in Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Theorem-of-Thought improves LLM reasoning accuracy on WebOfLies and MultiArith by selecting among three differently styled reasoning chains via NLI-based coherence scoring.

  2. Efficient Latent Semantic Clustering for Scaling Test-Time Computation of LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Latent Semantic Clustering uses a generator LLM's hidden states to group semantically equivalent outputs, removing the need for external NLI or embedding models in test-time scaling.

  3. Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes

    cs.CV 2025-12 conditional novelty 5.0 of 10

    SGPU trains a Gaussian process classifier on the eigenvalue spectrum of answer-embedding Gram matrices to estimate semantic uncertainty in LVLMs without clustering.

Pith tools