Pith. sign in

REVIEW 2 cited by

Spanish Pre-trained BERT Model and Evaluation Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.02976 v1 pith:R7V5KR2I submitted 2023-08-06 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords spanishlanguagemodelpre-traineddatabert-basedmodelstasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The Spanish language is one of the top 5 spoken languages in the world. Nevertheless, finding resources to train or evaluate Spanish language models is not an easy task. In this paper we help bridge this gap by presenting a BERT-based language model pre-trained exclusively on Spanish data. As a second contribution, we also compiled several tasks specifically for the Spanish language in a single repository much in the spirit of the GLUE benchmark. By fine-tuning our pre-trained Spanish model, we obtain better results compared to other BERT-based models pre-trained on multilingual corpora for most of the tasks, even achieving a new state-of-the-art on some of them. We have publicly released our model, the pre-training data, and the compilation of the Spanish benchmarks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 337 citations worldwide. Full citation record

  1. Examining Spanish Counseling with MIDAS: a Motivational Interviewing Dataset in Spanish

    cs.CL 2025-02 conditional novelty 6.0 of 10

    MIDAS, a new expert-annotated Spanish motivational interviewing dataset, reveals language-specific counselor behaviors and supports Spanish-language behavior classification.

  2. Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages

    cs.CL 2025-06

Pith tools