Pith. sign in

REVIEW 9 cited by

API Is Enough: Conformal Prediction for Large Language Models Without Logit-Access

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.01216 v2 pith:XCQ6AHX2 submitted 2024-03-02 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords llmsapproachlogit-accesspredictionwithoutapi-onlyconformalknown
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study aims to address the pervasive challenge of quantifying uncertainty in large language models (LLMs) without logit-access. Conformal Prediction (CP), known for its model-agnostic and distribution-free features, is a desired approach for various LLMs and data distributions. However, existing CP methods for LLMs typically assume access to the logits, which are unavailable for some API-only LLMs. In addition, logits are known to be miscalibrated, potentially leading to degraded CP performance. To tackle these challenges, we introduce a novel CP method that (1) is tailored for API-only LLMs without logit-access; (2) minimizes the size of prediction sets; and (3) ensures a statistical guarantee of the user-defined coverage. The core idea of this approach is to formulate nonconformity measures using both coarse-grained (i.e., sample frequency) and fine-grained uncertainty notions (e.g., semantic similarity). Experimental results on both close-ended and open-ended Question Answering tasks show our approach can mostly outperform the logit-based CP baselines.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Remember with Confidence: Uncertainty Quantification for Spatio-temporal Memory with Probabilistic Guarantees

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    Introduces object-level semantic uncertainty for VLM memory, the UQ-DAAAM refinement system, and probabilistic guarantees that selected high-quality views reduce uncertainty more effectively.

  2. Improving Backward Conformal Prediction via Non-Conformity Score Transformation

    stat.ML 2026-02 conditional novelty 7.0 of 10

    ST-BCP tightens the coverage bound in Backward Conformal Prediction by applying a computable data-dependent transformation to nonconformity scores, reducing the average gap from 4.20% to 1.12% on benchmarks while prov...

  3. Conformal Path Reasoning: Trustworthy Knowledge Graph Question Answering via Path-Level Calibration

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    CPR improves empirical coverage rate by 34% and reduces average prediction set size by 40% in KGQA benchmarks via query-level path calibration and RCVNet for discriminative nonconformity scores.

  4. Conformal Path Reasoning: Trustworthy Knowledge Graph Question Answering via Path-Level Calibration

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    CPR uses query-level conformal calibration over path scores and a PUCT-trained RCVNet to achieve valid coverage guarantees and smaller prediction sets in KGQA, reporting 45% higher empirical coverage and 52% smaller s...

  5. From Debate to Decision: Conformal Social Choice for Safe Multi-Agent Deliberation

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    Conformal Social Choice aggregates verbalized probabilities from LLM debates via linear opinion pooling and uses split conformal prediction to generate prediction sets that guarantee inclusion of the correct answer wi...

  6. Improving Backward Conformal Prediction via Non-Conformity Score Transformation

    stat.ML 2026-02 reject novelty 6.0 of 10

    A step-function score transformation I_w = w·1{s≥w} reduces the estimated-coverage gap in Backward Conformal Prediction from about 4.2% to 1.1% on CIFAR-10, CIFAR-100, and Tiny-ImageNet.

  7. Strategic Decision Support for AI Agents

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    The paper introduces an optimization framework for AI agents to strategically seek support, proving a threshold policy on support value and providing an online algorithm to control missed-support error without distrib...

  8. Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering

    cs.CL 2026-05 unverdicted novelty 5.0 of 10

    Mainstream UQ for LLMs reduces to unsupervised clustering of internal generation consistency and therefore cannot detect confident hallucinations or provide reliable safety signals.

  9. Differentiable Conformal Training for LLM Reasoning Factuality

    cs.LG 2026-04 unverdicted novelty 5.0 of 10

    DCF relaxes non-differentiable conformal factuality for LLM reasoning chains into a trainable form, yielding up to 141% higher retention of true claims on benchmarks while preserving reliability guarantees.

Pith tools