Pith. sign in

REVIEW 1 cited by

Pareto Optimal Learning for Estimating Large Language Model Errors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.16564 v4 pith:DWHCR4ZC submitted 2023-06-28 cs.CL stat.ML

classification cs.CLstat.ML
keywords errorinformationmethodmodelsparetolanguagelargeoptimal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have shown impressive abilities in many applications. When a concrete and precise answer is desired, it is important to have a quantitative estimation of the potential error rate. However, this can be challenging due to the text-in-text-out nature of generative models. We present a method based on Pareto optimization that generates a risk score to estimate the probability of error in an LLM response by integrating multiple sources of information. We prove theoretically that the error estimator optimized in our framework aligns with the LLM and the information sources in an Pareto optimal manner. Experimental results show that the risk scores estimated by our method are well correlated with the true LLM error rate, thus facilitating error correction. By dynamically combining with prompting strategies such as self-verification and information retrieval, we demonstrate the proposed method can be utilized to increase the performance of an LLM, surpassing state-of-the-art task specific models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts

    cs.CL 2024-12 accept novelty 6.0 of 10

    A personalized calibration network that combines LLM answers to multiple rubric questions predicted human judges' overall satisfaction scores on dialogues about twice as accurately as the uncalibrated LLM.

Pith tools