Pith. sign in

REVIEW 3 cited by

Tucano: Advancing Neural Text Generation for Portuguese

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.07854 v1 pith:YA4NYLY2 submitted 2024-11-12 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords portugueselanguagemodelsdevelopmentresourcestexttucanoavailable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Significant advances have been made in natural language processing in recent years. However, our current deep learning approach to language modeling requires substantial resources in terms of data and computation. One of the side effects of this data-hungry paradigm is the current schism between languages, separating those considered high-resource, where most of the development happens and resources are available, and the low-resource ones, which struggle to attain the same level of performance and autonomy. This study aims to introduce a new set of resources to stimulate the future development of neural text generation in Portuguese. In this work, we document the development of GigaVerbo, a concatenation of deduplicated Portuguese text corpora amounting to 200 billion tokens. Via this corpus, we trained a series of decoder-transformers named Tucano. Our models perform equal or superior to other Portuguese and multilingual language models of similar size in several Portuguese benchmarks. The evaluation of our models also reveals that model performance on many currently available benchmarks used by the Portuguese NLP community has little to no correlation with the scaling of token ingestion during training, highlighting the limitations of such evaluations when it comes to the assessment of Portuguese generative language models. All derivatives of our study are openly released on GitHub and Hugging Face. See https://nkluge-correa.github.io/Tucano/

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Study

    cs.CY 2025-12 conditional novelty 5.0 of 10

    In interviews with 11 Portuguese-language model developers, four AI ethics tools guided general ethical reflection but failed to surface Portuguese-specific harms like cultural misrepresentation and low language performance.

  2. BRoverbs -- Measuring how much LLMs understand Portuguese proverbs

    cs.CL 2025-09 conditional novelty 5.0 of 10

    BRoverbs lets researchers test whether language models understand Portuguese proverbs; commercial models nearly master it, small models often guess randomly.

  3. Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Study

    cs.CY 2025-12 unverdicted novelty 3.0 of 10

    Evaluation of Model Cards, ALTAI, FactSheets, and Harms Modeling on Portuguese language models shows they provide broad ethical guidance but overlook unique language features and negative impacts.

Pith tools