Pith. sign in

REVIEW 4 cited by

On Accurate Evaluation of GANs for Language Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1806.04936 v3 pith:O4LBEZQK submitted 2018-06-13 cs.CL

classification cs.CL
keywords languagemodelsgenerationbestbetterevaluationgansgenerated
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Generative Adversarial Networks (GANs) are a promising approach to language generation. The latest works introducing novel GAN models for language generation use n-gram based metrics for evaluation and only report single scores of the best run. In this paper, we argue that this often misrepresents the true picture and does not tell the full story, as GAN models can be extremely sensitive to the random initialization and small deviations from the best hyperparameter choice. In particular, we demonstrate that the previously used BLEU score is not sensitive to semantic deterioration of generated texts and propose alternative metrics that better capture the quality and diversity of the generated samples. We also conduct a set of experiments comparing a number of GAN models for text with a conventional Language Model (LM) and find that neither of the considered models performs convincingly better than the LM.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Autoregressive Text Generation Beyond Feedback Loops

    cs.LG 2019-08 conditional novelty 7.0 of 10

    A latent sequence model with a globally normalized pairwise CRF observation model generates coherent text while keeping state transitions non-autoregressive.

  2. ARAML: A Stable Adversarial Training Framework for Text Generation

    cs.CL 2019-08 conditional novelty 7.0 of 10

    ARAML stabilizes adversarial text generation by training the generator with reward-weighted maximum likelihood on samples drawn from a fixed distribution around real data instead of using policy gradient.

  3. Escaping the BLEU Trap: A Signal-Grounded Framework with Decoupled Semantic Guidance for EEG-to-Text Decoding

    cs.CL 2026-02 unverdicted novelty 6.0 of 10

    SemKey predicts four semantic attributes from EEG and conditions a frozen LLM on them, beating prior decoders on new semantic-alignment metrics while leaving true word-level accuracy low (2.7% content recall).

  4. IntentGPT: Few-shot Intent Discovery with Large Language Models

    cs.CL 2024-11 conditional novelty 6.0 of 10

    A training-free LLM prompting pipeline with semantic few-shot retrieval and feedback of discovered intents outperforms trained baselines on few-shot intent discovery benchmarks.

Pith tools