REVIEW 10 cited by
All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Human evaluations are typically considered the gold standard in natural language generation, but as models' fluency improves, how well can evaluators detect and judge machine-generated text? We run a study assessing non-experts' ability to distinguish between human- and machine-authored text (GPT2 and GPT3) in three domains (stories, news articles, and recipes). We find that, without training, evaluators distinguished between GPT3- and human-authored text at random chance level. We explore three approaches for quickly training evaluators to better identify GPT3-authored text (detailed instructions, annotated examples, and paired examples) and find that while evaluators' accuracy improved up to 55%, it did not significantly improve across the three domains. Given the inconsistent results across text domains and the often contradictory reasons evaluators gave for their judgments, we examine the role untrained human evaluations play in NLG evaluation and provide recommendations to NLG researchers for improving human evaluations of text generated from state-of-the-art models.
Forward citations
Cited by 10 Pith papers
-
Towards Efficient and Effective Alignment of Large Language Models
A thesis presenting Lion, WebR, LTE, BMC, and FollowBench, five empirical methods that together address LLM alignment data, training, and evaluation.
-
GaussMark: A Practical Approach for Structural Watermarking of Language Models
GaussMark embeds a detectable watermark by adding per-generation Gaussian noise to one weight matrix and detecting gradient alignment with that noise.
-
Towards Trustworthy AI: Characterizing User-Reported Risks across LLMs "In the Wild"
Across seven AI chatbots, Reddit users report mostly reliability failures, with each chatbot showing a distinct pattern of safety, privacy, and security complaints.
-
Is Your LLM-Based Multi-Agent a Reliable Real-World Planner? Exploring Fraud Detection in Travel Planning
Multi-agent LLM travel planners are frequently deceived by injected fake listings, coordinated fake reviews, and multi-round scam conversations, and a simple anti-fraud reviewer helps only some models.
-
Mind the Gap! Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different Languages
Users who first used a Spanish AI writing assistant subsequently used the English AI writing assistant less, suggesting a spillover that violates choice independence.
-
The unintended consequences of large language models as a labor-augmenting technology in science
LLM speed-ups in discovery or production raise the opportunity cost of researcher time, making scientists more selective in some phases, less thorough in most, and only deeper when follow-up work itself is accelerated.
-
Psychology-Driven Enhancement of Humour Translation
A decomposition-and-recomposition prompt method for humor translation reports large gains on LLM-based metrics, but the evaluation lacks human validation and statistical checks.
-
Using Machine Learning to Distinguish Human-written from Machine-generated Creative Fiction
Naive Bayes and MLP classifiers distinguish about 100-word excerpts of human-written detective fiction from ChatGPT-3.5 output with roughly 96% accuracy on in-domain tests.
-
LLM Encoder vs. Decoder: Robust Detection of Chinese AI-Generated Text with LoRA
On the NLPCC 2025 Chinese AI-text detection benchmark, LoRA-adapted Qwen2.5-7B reaches 95.94% test accuracy, beating BERT-large (79.3%), RoBERTa-large (76.3%), and FastText (83.5%).
-
The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
A review of known vulnerabilities in LLM benchmarks (contamination, overfitting, human and LLM judge bias) with a sketch of a proposed zero-day evaluation framework.
Discussion (0). Continue with ORCID to comment.