Pith. sign in

REVIEW 1 cited by

Quati: A Brazilian Portuguese Information Retrieval Dataset from Native Speakers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.06976 v1 pith:E7Y4FOIW submitted 2024-04-10 cs.IR

classification cs.IR
keywords quatidatasetsportuguesebrazilianhigh-qualitydatasetdocumentshttps
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite Portuguese being one of the most spoken languages in the world, there is a lack of high-quality information retrieval datasets in that language. We present Quati, a dataset specifically designed for the Brazilian Portuguese language. It comprises a collection of queries formulated by native speakers and a curated set of documents sourced from a selection of high-quality Brazilian Portuguese websites. These websites are frequented more likely by real users compared to those randomly scraped, ensuring a more representative and relevant corpus. To label the query-document pairs, we use a state-of-the-art LLM, which shows inter-annotator agreement levels comparable to human performance in our assessments. We provide a detailed description of our annotation methodology to enable others to create similar datasets for other languages, providing a cost-effective way of creating high-quality IR datasets with an arbitrary number of labeled documents per query. Finally, we evaluate a diverse range of open-source and commercial retrievers to serve as baseline systems. Quati is publicly available at https://huggingface.co/datasets/unicamp-dl/quati and all scripts at https://github.com/unicamp-dl/quati .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When LLMs Disagree: Diagnosing Relevance Filtering Bias and Retrieval Divergence in SDG Search

    cs.IR 2025-07 conditional novelty 4.0 of 10

    Two LLMs disagree on about 16% of SDG relevance labels, and the disagreement is lexically systematic and changes top-20 retrieval results.

Pith tools