Pith. sign in

REVIEW 4 cited by

ConU: Conformal Uncertainty in Large Language Models with Correctness Coverage Guarantees

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.00499 v3 pith:KD4JKXZR submitted 2024-06-29 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords uncertaintyconformalcorrectnesslanguagellmspredictioncoverageguarantees
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Uncertainty quantification (UQ) in natural language generation (NLG) tasks remains an open challenge, exacerbated by the closed-source nature of the latest large language models (LLMs). This study investigates applying conformal prediction (CP), which can transform any heuristic uncertainty notion into rigorous prediction sets, to black-box LLMs in open-ended NLG tasks. We introduce a novel uncertainty measure based on self-consistency theory, and then develop a conformal uncertainty criterion by integrating the uncertainty condition aligned with correctness into the CP algorithm. Empirical evaluations indicate that our uncertainty measure outperforms prior state-of-the-art methods. Furthermore, we achieve strict control over the correctness coverage rate utilizing 7 popular LLMs on 4 free-form NLG datasets, spanning general-purpose and medical scenarios. Additionally, the calibrated prediction sets with small size further highlights the efficiency of our method in providing trustworthy guarantees for practical open-ended NLG applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. I'm Spartacus, No, I'm Spartacus: Measuring and Understanding LLM Identity Confusion

    cs.CR 2024-11 reject novelty 5.0 of 10

    Seven of 27 tested LLMs (25.93%) exhibited identity confusion, which the authors link to hallucination and show reduces user trust, especially in critical tasks.

  2. A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A review that organizes LLM uncertainty quantification into token-level, self-verbalized, semantic-similarity, and mechanistic interpretability categories.

  3. AutoIoT: Automated IoT Platform Using Large Language Models

    cs.CR 2024-11 reject novelty 4.0 of 10

    AutoIoT generates conflict-free smart home automation rules from user photos and manuals using LLMs, and verifies them with Maude via a logic-syntax code adapter.

  4. Assessing GPT Model Uncertainty in Mathematical OCR Tasks via Entropy Analysis

    cs.IT 2024-12 reject novelty 3.0 of 10

    The paper reports that GPT-4o's token-level uncertainty, computed as the negative log-likelihood of its output, rises monotonically as image resolution falls from 300 to 72 dpi on a single test page.

Pith tools