Pith. sign in

Identifying Well-formed Natural Language Questions

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Understanding search queries is a hard problem as it involves dealing with "word salad" text ubiquitously issued by users. However, if a query resembles a well-formed question, a natural language processing pipeline is able to perform more accurate interpretation, thus reducing downstream compounding errors. Hence, identifying whether or not a query is well formed can enhance query understanding. Here, we introduce a new task of identifying a well-formed natural language question. We construct and release a dataset of 25,100 publicly available questions classified into well-formed and non-wellformed categories and report an accuracy of 70.7% on the test set. We also show that our classifier can be used to improve the performance of neural sequence-to-sequence models for generating questions for reading comprehension.

citation-role summary

dataset 1

citation-polarity summary

fields

cs.AI 1

years

2024 1

verdicts

CONDITIONAL 1

roles

dataset 1

polarities

use dataset 1

representative citing papers

Bridging the Data Provenance Gap Across Text, Speech and Video

cs.AI · 2024-12-19 · conditional · novelty 6.0

A manual audit of nearly 4,000 text, speech, and video datasets finds AI training data increasingly comes from web and social media sources, carries hidden non-commercial restrictions, and remains Western-centric with no relative diversity gains since 2013.

citing papers explorer

Showing 1 of 1 citing paper.

  • Bridging the Data Provenance Gap Across Text, Speech and Video cs.AI · 2024-12-19 · conditional · none · ref 250 · internal anchor

    A manual audit of nearly 4,000 text, speech, and video datasets finds AI training data increasingly comes from web and social media sources, carries hidden non-commercial restrictions, and remains Western-centric with no relative diversity gains since 2013.