REVIEW 4 cited by
On Accurate Evaluation of GANs for Language Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Generative Adversarial Networks (GANs) are a promising approach to language generation. The latest works introducing novel GAN models for language generation use n-gram based metrics for evaluation and only report single scores of the best run. In this paper, we argue that this often misrepresents the true picture and does not tell the full story, as GAN models can be extremely sensitive to the random initialization and small deviations from the best hyperparameter choice. In particular, we demonstrate that the previously used BLEU score is not sensitive to semantic deterioration of generated texts and propose alternative metrics that better capture the quality and diversity of the generated samples. We also conduct a set of experiments comparing a number of GAN models for text with a conventional Language Model (LM) and find that neither of the considered models performs convincingly better than the LM.
Forward citations
Cited by 4 Pith papers
-
Autoregressive Text Generation Beyond Feedback Loops
A latent sequence model with a globally normalized pairwise CRF observation model generates coherent text while keeping state transitions non-autoregressive.
-
ARAML: A Stable Adversarial Training Framework for Text Generation
ARAML stabilizes adversarial text generation by training the generator with reward-weighted maximum likelihood on samples drawn from a fixed distribution around real data instead of using policy gradient.
-
Escaping the BLEU Trap: A Signal-Grounded Framework with Decoupled Semantic Guidance for EEG-to-Text Decoding
SemKey predicts four semantic attributes from EEG and conditions a frozen LLM on them, beating prior decoders on new semantic-alignment metrics while leaving true word-level accuracy low (2.7% content recall).
-
IntentGPT: Few-shot Intent Discovery with Large Language Models
A training-free LLM prompting pipeline with semantic few-shot retrieval and feedback of discovered intents outperforms trained baselines on few-shot intent discovery benchmarks.
Discussion (0). Continue with ORCID to comment.