deltaBLEU: A Discriminative Metric for Generation Tasks with Intrinsically Diverse Targets

Alessandro Sordoni; Bill Dolan; Chris Brockett; Chris Quirk; Jianfeng Gao; Margaret Mitchell; Michael Auli; Michel Galley; Yangfeng Ji

arxiv: 1506.06863 · v2 · pith:CYCURKQTnew · submitted 2015-06-23 · 💻 cs.CL

deltaBLEU: A Discriminative Metric for Generation Tasks with Intrinsically Diverse Targets

Michel Galley , Chris Brockett , Alessandro Sordoni , Yangfeng Ji , Michael Auli , Chris Quirk , Margaret Mitchell , Jianfeng Gao

show 1 more author

Bill Dolan

This is my paper

classification 💻 cs.CL

keywords bleudeltableutasksdiscriminativediversegenerationhumanmetric

0 comments

read the original abstract

We introduce Discriminative BLEU (deltaBLEU), a novel metric for intrinsic evaluation of generated text in tasks that admit a diverse range of possible outputs. Reference strings are scored for quality by human raters on a scale of [-1, +1] to weight multi-reference BLEU. In tasks involving generation of conversational responses, deltaBLEU correlates reasonably with human judgments and outperforms sentence-level and IBM BLEU in terms of both Spearman's rho and Kendall's tau.

This paper has not been read by Pith yet.

deltaBLEU: A Discriminative Metric for Generation Tasks with Intrinsically Diverse Targets

discussion (0)