Pith. sign in

REVIEW 4 cited by

Perplexity from PLM Is Unreliable for Evaluating Text Quality

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.05892 v2 pith:BFHNBXRD submitted 2022-10-12 cs.CL cs.AI

classification cs.CLcs.AI
keywords textqualityevaluatingevaluategeneratedperformanceperplexityunreliable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, amounts of works utilize perplexity~(PPL) to evaluate the quality of the generated text. They suppose that if the value of PPL is smaller, the quality(i.e. fluency) of the text to be evaluated is better. However, we find that the PPL referee is unqualified and it cannot evaluate the generated text fairly for the following reasons: (i) The PPL of short text is larger than long text, which goes against common sense, (ii) The repeated text span could damage the performance of PPL, and (iii) The punctuation marks could affect the performance of PPL heavily. Experiments show that the PPL is unreliable for evaluating the quality of given text. Last, we discuss the key problems with evaluating text quality using language models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BenCzechMark : A Czech-centric Multitask and Multimetric Benchmark for Large Language Models with Duel Scoring Mechanism

    cs.CL 2024-12 conditional novelty 7.0 of 10

    BenCzechMark is a new 50-task Czech benchmark with a duel scoring system, a 320GB Czech corpus, and a leaderboard of 50 open-weight models.

  2. FitCF: A Framework for Automatic Feature Importance-guided Counterfactual Example Generation

    cs.CL 2025-01 conditional novelty 6.0 of 10

    ZeroCF and FitCF generate label-flipping text counterfactuals from BERT feature attributions, and FitCF outperforms Polyjuice, BAE, and FIZLE on AG News and SST2.

  3. How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives

    cs.CL 2024-12 conditional novelty 6.0 of 10

    This paper introduces an automated evaluation framework with extraction-based faithfulness metrics, perplexity for assumptions, and embedding-based human similarity, and shows it can reveal LLM sign self-correction on...

  4. ASRank: Zero-Shot Re-Ranking with Answer Scent for Document Retrieval

    cs.CL 2025-01 conditional novelty 5.0 of 10

    ASRank re-ranks retrieved documents by scoring how well each document supports a zero-shot answer scent generated by a large LLM, beating UPR and RankGPT on several QA datasets.

Pith tools