Pith. sign in

REVIEW 1 cited by

A Chinese Multi-type Complex Questions Answering Dataset over Wikidata

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.06086 v1 pith:PQTGRQEV submitted 2021-11-11 cs.CL cs.AIcs.DB

classification cs.CLcs.AIcs.DB
keywords questionscomplexdatasetwikidatachinesekbqaknowledgeanswering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Complex Knowledge Base Question Answering is a popular area of research in the past decade. Recent public datasets have led to encouraging results in this field, but are mostly limited to English and only involve a small number of question types and relations, hindering research in more realistic settings and in languages other than English. In addition, few state-of-the-art KBQA models are trained on Wikidata, one of the most popular real-world knowledge bases. We propose CLC-QuAD, the first large scale complex Chinese semantic parsing dataset over Wikidata to address these challenges. Together with the dataset, we present a text-to-SPARQL baseline model, which can effectively answer multi-type complex questions, such as factual questions, dual intent questions, boolean questions, and counting questions, with Wikidata as the background knowledge. We finally analyze the performance of SOTA KBQA models on this dataset and identify the challenges facing Chinese KBQA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Conversational Lexicography: Querying Lexicographic Data on Knowledge Graphs with SPARQL through Natural Language

    cs.CL 2025-05 conditional novelty 6.0 of 10

    This paper introduces a taxonomy and 1.27M-example template dataset for text-to-SPARQL on Wikidata lexicographic data, and finds that only GPT-3.5-Turbo generalizes to novel query types.

Pith tools