Pith. sign in

SINA-BERT: A pre-trained Language Model for Analysis of Medical Texts in Persian

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We have released Sina-BERT, a language model pre-trained on BERT (Devlin et al., 2018) to address the lack of a high-quality Persian language model in the medical domain. SINA-BERT utilizes pre-training on a large-scale corpus of medical contents including formal and informal texts collected from a variety of online resources in order to improve the performance on health-care related tasks. We employ SINA-BERT to complete following representative tasks: categorization of medical questions, medical sentiment analysis, and medical question retrieval. For each task, we have developed Persian annotated data sets for training and evaluation and learnt a representation for the data of each task especially complex and long medical questions. With the same architecture being used across tasks, SINA-BERT outperforms BERT-based models that were previously made available in the Persian language.

citation-role summary

background 1

citation-polarity summary

fields

cs.CL 1

years

2026 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

Gaokerena: A Small Persian Medical Language Model Family

cs.CL · 2026-08-02 · conditional · novelty 5.0

Fine-tuned Persian medical language models reach 49-53% on translated medical MMLU, with datasets released, but the reasoning variant's gain depends on extra test-time compute and a verifier.

citing papers explorer

Showing 1 of 1 citing paper.

  • Gaokerena: A Small Persian Medical Language Model Family cs.CL · 2026-08-02 · conditional · none · ref 3 · internal anchor

    Fine-tuned Persian medical language models reach 49-53% on translated medical MMLU, with datasets released, but the reasoning variant's gain depends on extra test-time compute and a verifier.