Pith. sign in

REVIEW 3 cited by

Unsupervised Evaluation of Interactive Dialog with DialoGPT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.12719 v1 pith:5HQAJLEM submitted 2020-06-23 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords dialogevaluationfine-grainedmetricautomaticdialogptintroduceslevels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

It is important to define meaningful and interpretable automatic evaluation metrics for open-domain dialog research. Standard language generation metrics have been shown to be ineffective for dialog. This paper introduces the FED metric (fine-grained evaluation of dialog), an automatic evaluation metric which uses DialoGPT, without any fine-tuning or supervision. It also introduces the FED dataset which is constructed by annotating a set of human-system and human-human conversations with eighteen fine-grained dialog qualities. The FED metric (1) does not rely on a ground-truth response, (2) does not require training data and (3) measures fine-grained dialog qualities at both the turn and whole dialog levels. FED attains moderate to strong correlation with human judgement at both levels.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes

    cs.AI 2026-07 conditional novelty 6.0 of 10

    MAR-12 improves humor and hate detection in memes by prompting a VLM through twelve reasoning perspectives, attention-weighting them, and generating explanations from the weighted evidence.

  2. Multilingual Prompt Engineering in Large Language Models: A Survey Across NLP Tasks

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A survey that categorizes multilingual prompting techniques by NLP task and language family, and designates potential state-of-the-art prompting methods for each dataset.

  3. An Explainable Emotion Alignment Framework for LLM-Empowered Agent in Metaverse Service Ecosystem

    cs.AI 2025-07 reject novelty 3.0 of 10

    Adding hand-coded PAD emotions and chain-of-thought rationales to LLM agents made their simulated delivery behavior look more like real rider data, though no quantitative test confirms it.

Pith tools