Pith. sign in

REVIEW 2 cited by

Trading Off Diversity and Quality in Natural Language Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.10450 v1 pith:7MRUSENK submitted 2020-04-22 cs.CL

classification cs.CL
keywords decodingqualitydiversitygenerationsamplingalgorithmlanguagelikelihood
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

For open-ended language generation tasks such as storytelling and dialogue, choosing the right decoding algorithm is critical to controlling the tradeoff between generation quality and diversity. However, there presently exists no consensus on which decoding procedure is best or even the criteria by which to compare them. We address these issues by casting decoding as a multi-objective optimization problem aiming to simultaneously maximize both response quality and diversity. Our framework enables us to perform the first large-scale evaluation of decoding methods along the entire quality-diversity spectrum. We find that when diversity is a priority, all methods perform similarly, but when quality is viewed as more important, the recently proposed nucleus sampling (Holtzman et al. 2019) outperforms all other evaluated decoding algorithms. Our experiments also confirm the existence of the `likelihood trap', the counter-intuitive observation that high likelihood sequences are often surprisingly low quality. We leverage our findings to create and evaluate an algorithm called \emph{selective sampling} which tractably approximates globally-normalized temperature sampling.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Comparative Study of Decoding Strategies in Medical Text Generation

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Across five medical text tasks, deterministic decoding methods generally score higher than stochastic sampling, while medical-specific models do not outperform general models and are more sensitive to decoding choice.

  2. Breaking the Likelihood Trap: Consistent Generative Recommendation with Graph-structured Model

    cs.IR 2025-10 conditional novelty 4.0 of 10

    CONGRATS uses a DAG-structured positional decoder and evaluator-in-the-loop training to generate more diverse and accurate recommendation lists, showing offline and Kuaishou A/B gains.

Pith tools