Pith. sign in

Eval all, trust a few, do wrong to none: Comparing sentence generation models

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

In this paper, we study recent neural generative models for text generation related to variational autoencoders. Previous works have employed various techniques to control the prior distribution of the latent codes in these models, which is important for sampling performance, but little attention has been paid to reconstruction error. In our study, we follow a rigorous evaluation protocol using a large set of previously used and novel automatic and human evaluation metrics, applied to both generated samples and reconstructions. We hope that it will become the new evaluation standard when comparing neural generative models for text.

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Measuring Diversity in Synthetic Datasets

cs.CL · 2025-02-12 · conditional · novelty 6.0

DCScore measures dataset diversity as the sum of self-classification probabilities under a softmax similarity matrix, and the paper shows it tracks generation temperature, human judgment, and LLM rankings.

citing papers explorer

Showing 1 of 1 citing paper.

  • Measuring Diversity in Synthetic Datasets cs.CL · 2025-02-12 · conditional · none · ref 9 · internal anchor

    DCScore measures dataset diversity as the sum of self-classification probabilities under a softmax similarity matrix, and the paper shows it tracks generation temperature, human judgment, and LLM rankings.