REVIEW 2 cited by
AdTEC: A Unified Benchmark for Evaluating Text Quality in Search Engine Advertising
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the increase in the fluency of ad texts automatically created by natural language generation technology, there is high demand to verify the quality of these creatives in a real-world setting. We propose AdTEC (Ad Text Evaluation Benchmark by CyberAgent), the first public benchmark to evaluate ad texts from multiple perspectives within practical advertising operations. Our contributions are as follows: (i) Defining five tasks for evaluating the quality of ad texts, as well as building a Japanese dataset based on the practical operational experiences of building a Japanese dataset based on the practical operational experiences of advertising agencies, which are typically kept in-house. (ii) Validating the performance of existing pre-trained language models (PLMs) and human evaluators on the dataset. (iii) Analyzing the characteristics and providing challenges of the benchmark. The results show that while PLMs have already reached practical usage level in several tasks, humans still outperform in certain domains, implying that there is significant room for improvement in this area.
Forward citations
Cited by 2 Pith papers
-
AdParaphrase: Paraphrase Dataset for Analyzing Linguistic Features toward Generating Attractive Ad Texts
AdParaphrase provides 725 human-preference-annotated paraphrase pairs of Japanese ad texts and shows fluency, length, noun count, and bracket use correlate with attractiveness.
-
Exploring the Relationship Between Diversity and Quality in Ad Text Generation
In Japanese ad text generation, methods that increase output diversity generally lower ad quality, with beam search and sampling responding differently to examples and output counts.
Discussion (0). Continue with ORCID to comment.