REVIEW 1 cited by
Open Sentence Embeddings for Portuguese with the Serafim PT* encoders family
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Open Sentence Embeddings for Portuguese with the Serafim PT* encoders family
read the original abstract
Sentence encoder encode the semantics of their input, enabling key downstream applications such as classification, clustering, or retrieval. In this paper, we present Serafim PT*, a family of open-source sentence encoders for Portuguese with various sizes, suited to different hardware/compute budgets. Each model exhibits state-of-the-art performance and is made openly available under a permissive license, allowing its use for both commercial and research purposes. Besides the sentence encoders, this paper contributes a systematic study and lessons learned concerning the selection criteria of learning objectives and parameters that support top-performing encoders.
Forward citations
Cited by 1 Pith paper
-
MTEB-BR: A Text Embedding Benchmark for Brazilian Portuguese
A native 22-task Brazilian-Portuguese embedding benchmark cleanly tiers 93 models, places an open model in the unresolved top tier, and finds only moderate rank correlation (ρ=0.75) with the multilingual MTEB board.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.