Pith. sign in

REVIEW 2 cited by

Advancing Neural Encoding of Portuguese with Transformer Albertina PT-*

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.06721 v2 pith:MENTTJRE submitted 2023-05-11 cs.CL

classification cs.CL
keywords portuguesealbertinapt-brlanguagept-ptsetsdataencoding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To advance the neural encoding of Portuguese (PT), and a fortiori the technological preparation of this language for the digital age, we developed a Transformer-based foundation model that sets a new state of the art in this respect for two of its variants, namely European Portuguese from Portugal (PT-PT) and American Portuguese from Brazil (PT-BR). To develop this encoder, which we named Albertina PT-*, a strong model was used as a starting point, DeBERTa, and its pre-training was done over data sets of Portuguese, namely over data sets we gathered for PT-PT and PT-BR, and over the brWaC corpus for PT-BR. The performance of Albertina and competing models was assessed by evaluating them on prominent downstream language processing tasks adapted for Portuguese. Both Albertina PT-PT and PT-BR versions are distributed free of charge and under the most permissive license possible and can be run on consumer-grade hardware, thus seeking to contribute to the advancement of research and innovation in language technology for Portuguese.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus

    cs.CL 2026-04 unverdicted novelty 7.0 of 10

    NorBERTo, a ModernBERT-style Portuguese encoder trained from scratch on the 331B-token Aurora-PT corpus, posts top scores on PLUE and ASSIN 2 entailment, but lower scores on semantic similarity.

  2. Uncovering Latent Connections in Indigenous Heritage: Semantic Pipelines for Cultural Preservation in Brazil

    cs.HC 2025-07 conditional novelty 4.0 of 10

    The paper introduces two embedding pipelines and a visualization tool that reveal latent clusters and label errors in the Museu Nacional dos Povos Indígenas digital collection.

Pith tools