REVIEW 10 cited by
Calibration of Encoder Decoder Models for Neural Machine Translation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We study the calibration of several state of the art neural machine translation(NMT) systems built on attention-based encoder-decoder models. For structured outputs like in NMT, calibration is important not just for reliable confidence with predictions, but also for proper functioning of beam-search inference. We show that most modern NMT models are surprisingly miscalibrated even when conditioned on the true previous tokens. Our investigation leads to two main reasons -- severe miscalibration of EOS (end of sequence marker) and suppression of attention uncertainty. We design recalibration methods based on these signals and demonstrate improved accuracy, better sequence-level calibration, and more intuitive results from beam-search.
Forward citations
Cited by 10 Pith papers
-
ECUAS$_n$: A family of metrics for principled evaluation of uncertainty-augmented systems
Proposes ECUAS_n metrics as proper scoring rules for evaluating uncertainty-augmented systems, with n controlling cost trade-offs between predictions and uncertainties.
-
On NMT Search Errors and Model Errors: Cat Got Your Tongue?
Beam search misses the global best translation for over half of sentences; under exact search, neural MT models prefer the empty translation for more than 50% of inputs.
-
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
A classifier trained on layerwise decodes and BERTScore similarities to the final output gives better-calibrated confidence for tool calls and improves expected utility at medium and high risk levels.
-
Calibrating Translation Decoding with Quality Estimation on LLMs
Optimizing the Pearson correlation between hypothesis likelihood and an external quality score during fine-tuning improves LLM translation quality and turns log-likelihood into a competitive reference-free quality estimator.
-
Optimizing Temperature for Language Models with Multi-Sample Inference
Selecting the temperature at the entropy turning point of a language model's generated text yields near-optimal multi-sample inference accuracy without labeled validation data.
-
Variability Need Not Imply Error: The Case of Adequate but Semantically Distinct Responses
PROBAR, the estimated probability that a language model's sampled responses are adequate to the prompt, outperforms semantic entropy for selective prediction across ambiguous and open-ended prompts.
-
Adaptive Decoding via Latent Preference Optimization
Adaptive Decoding learns to pick a discrete sampling temperature at token or sequence level via a DPO-style loss over latent temperature choices, and beats fixed temperatures on average across three task families.
-
Speaking in Self-Assessing Tongues: On the Verbalized Confidence of LLMs in Machine Translation
Empirical study finds verbalized per-token confidence methods in LLMs for MT perform similarly to internal signals on error detection and calibration but show little correlation.
-
Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics
Two probability-only confidence metrics, a top-to-kth beam ratio and a tail-thinness score, improve quality correlation for BART and Flan-T5 on several summarization, translation, and QA datasets.
-
Text-to-SQL Calibration: No Need to Ask -- Just Rescale Model Probabilities
Product-of-token-probabilities with Platt or isotonic rescaling is a strong, cheap confidence signal for text-to-SQL, beating minimum-token pooling but matching or losing to self-check on large Llama models.
Discussion (0). Continue with ORCID to comment.