REVIEW 21 cited by
Calibrate Before Use: Improving Few-Shot Performance of Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
GPT-3 can perform numerous tasks when provided a natural language prompt that contains a few training examples. We show that this type of few-shot learning can be unstable: the choice of prompt format, training examples, and even the order of the training examples can cause accuracy to vary from near chance to near state-of-the-art. We demonstrate that this instability arises from the bias of language models towards predicting certain answers, e.g., those that are placed near the end of the prompt or are common in the pre-training data. To mitigate this, we first estimate the model's bias towards each answer by asking for its prediction when given the training prompt and a content-free test input such as "N/A". We then fit calibration parameters that cause the prediction for this input to be uniform across answers. On a diverse set of tasks, this contextual calibration procedure substantially improves GPT-3 and GPT-2's average accuracy (up to 30.0% absolute) and reduces variance across different choices of the prompt.
Forward citations
Cited by 21 Pith papers
-
EsBBQ and CaBBQ: The Spanish and Catalan Bias Benchmarks for Question Answering
EsBBQ and CaBBQ are new Spanish and Catalan bias benchmarks for multiple-choice QA, built with survey-validated stereotypes from Spain and evaluated on 17 language models.
-
Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting
An early-exit rule with a zero-shot fallback, calibrated by Learn-then-Test risk control, keeps the average loss from corrupted in-context demonstrations under a preset bound.
-
MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations
MedBench-IT collects 17,410 Italian medical entrance exam questions and reports model accuracy, response consistency, ordering bias, reasoning-prompt effects, and readability correlations.
-
Characterizing Fitness Landscape Structures in Prompt Engineering
Prompt fitness autocorrelation appears smooth under systematic enumeration but rugged with an intermediate-distance peak under novelty-driven sampling, yet the two analyses cover non-overlapping distance ranges.
-
MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers
A zero-shot mixture of sparse, dense, and simulated human retrievers, weighted by pre- and post-retrieval geometry signals, beats individual small retrievers and 7B LLM retrievers on four scientific retrieval benchmarks.
-
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
A classifier trained on layerwise decodes and BERTScore similarities to the final output gives better-calibrated confidence for tool calls and improves expected utility at medium and high risk levels.
-
Time Will Tell: Timing Side Channels via Output Token Count in Large Language Models
Output token count, observable through response timing, can reveal a user's target language or classification result with 70-87% accuracy in the authors' experiments.
-
Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages
A translated temporal reasoning dataset plus a cross-lingual example retriever that outperforms semantic-alignment baselines on low-resource language temporal questions.
-
AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models
AURA uses step-level reward models, self-critique, and safety-aware decoding to reduce affordance-based safety failures in LLM outputs.
-
StaICC: Standardized Evaluation for Classification Task in In-context Learning
StaICC standardizes in-context classification evaluation with fixed prompts and splits, then measures 29 LMs and 10 inference methods under those fixed settings.
-
From Words to Workflows: Automating Business Processes
Text2Workflow is a multi-prompt LLM system with human feedback that generates JSON workflows from natural language, scoring 71.3% average semantic accuracy on the authors' 60-request Process2JSON dataset, versus 64.2%...
-
Does Prompt Formatting Have Any Impact on LLM Performance?
Prompt format alone can shift GPT model scores by double-digit percentage points, with GPT-3.5 more affected than GPT-4 and no universally best template.
-
(Towards) Scalable Reliable Automated Evaluation with Large Language Models
Multi-LLM pairwise Elo ranking with adjustable consensus thresholds produces rankings of competency profiles that average Spearman ρ≈0.83 with expert judgments.
-
Fine-tuning on simulated data outperforms prompting for agent tone of voice
Fine-tuning a 1B-parameter LLM on as few as 100 synthetically generated, readability-filtered samples achieved conversational tone more reliably than a verbose system prompt.
-
Dynamic Context-Aware Prompt Recommendation for Domain-Specific AI Applications
A dynamic prompt recommendation system for skill-based security copilots combines retrieval, hierarchical skill selection, and telemetry-based ranking, reporting high usefulness in internal evaluations.
-
Extending Epistemic Uncertainty Beyond Parameters Would Assist in Designing Reliable LLMs
Bayesian Modeling of Experiments is proposed as a unifying framework for quantifying and reducing the many sources of uncertainty in LLM deployments, beyond abstention.
-
Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods
LLMs extend, rather than replace, classical social science methods, with a proposed three-tier bias framework for LLM-augmented surveys.
-
Few-Shot Optimization for Sensor Data Using Large Language Models: A Case Study on Fatigue Detection
A hybrid Euclidean-distance and LLM-relevance example selector for few-shot sensor classification reports a small, statistically fragile gain over distance-only selection on a fatigue detection dataset.
-
Leveraging Large Language Models for enzymatic reaction prediction and characterization
Fine-tuned Llama-3.1 models can perform enzymatic reaction prediction tasks, and multitask learning improves forward and retrosynthesis over single-task training.
-
LLM Ethics Benchmark: A Three-Dimensional Assessment System for Evaluating Moral Reasoning in Large Language Models
A proposed three-dimensional benchmark for LLM moral reasoning that combines MFQ, WVS, and moral dilemmas, but the reported model scores are not reproducible from the paper.
-
Efficient Knowledge Feeding to Language Models: A Novel Integrated Encoder-Decoder Architecture
A retrieval-augmented encoder-decoder that injects 'in-context vectors' into latent states is presented, with claims of competing with much larger RAG models on three QA benchmarks.
Discussion (0). Continue with ORCID to comment.