Pith. sign in

VinaLLaMA: LLaMA-based Vietnamese Foundation Model

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

In this technical report, we present VinaLLaMA, an open-weight, state-of-the-art (SOTA) Large Language Model for the Vietnamese language, built upon LLaMA-2 with an additional 800 billion trained tokens. VinaLLaMA not only demonstrates fluency in Vietnamese but also exhibits a profound understanding of Vietnamese culture, making it a truly indigenous model. VinaLLaMA-7B-chat, trained on 1 million high-quality synthetic samples, achieves SOTA results on key benchmarks, including VLSP, VMLU, and Vicuna Benchmark Vietnamese, marking a significant advancement in the Vietnamese AI landscape and offering a versatile resource for various applications.

fields

cs.AI 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

Soro: A Lightweight Foundation Model and Chatbot for Tajik

cs.AI · 2026-04-09 · conditional · novelty 5.5

Tajik-specialized Gemma 3 derivatives (12B/27B) beat same-size baselines by ~6–8 points on new Tajik exams after 1.9B-token continual pretraining, with FP8/INT4 still usable on edge GPUs.

citing papers explorer

Showing 1 of 1 citing paper.

  • Soro: A Lightweight Foundation Model and Chatbot for Tajik cs.AI · 2026-04-09 · conditional · none · ref 3 · internal anchor

    Tajik-specialized Gemma 3 derivatives (12B/27B) beat same-size baselines by ~6–8 points on new Tajik exams after 1.9B-token continual pretraining, with FP8/INT4 still usable on edge GPUs.