REVIEW 3 cited by
VinaLLaMA: LLaMA-based Vietnamese Foundation Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this technical report, we present VinaLLaMA, an open-weight, state-of-the-art (SOTA) Large Language Model for the Vietnamese language, built upon LLaMA-2 with an additional 800 billion trained tokens. VinaLLaMA not only demonstrates fluency in Vietnamese but also exhibits a profound understanding of Vietnamese culture, making it a truly indigenous model. VinaLLaMA-7B-chat, trained on 1 million high-quality synthetic samples, achieves SOTA results on key benchmarks, including VLSP, VMLU, and Vicuna Benchmark Vietnamese, marking a significant advancement in the Vietnamese AI landscape and offering a versatile resource for various applications.
Forward citations
Cited by 3 Pith papers
-
Soro: A Lightweight Foundation Model and Chatbot for Tajik
Tajik-specialized Gemma 3 derivatives (12B/27B) beat same-size baselines by ~6–8 points on new Tajik exams after 1.9B-token continual pretraining, with FP8/INT4 still usable on edge GPUs.
-
VietMed-MCQ: A Consistency-Filtered Data Synthesis Framework for Vietnamese Traditional Medicine Evaluation
A dual-model consistency filter and a 94.2% expert-approved sample back VietMed-MCQ, a 3,190-question Vietnamese Traditional Medicine quiz where Qwen2.5 models beat Vietnamese-specialized models by 7.2%.
-
ViBidirectionMT-Eval: Machine Translation for Vietnamese-Chinese and Vietnamese-Lao language pair
A shared-task report and dataset release for four Vietnamese-Chinese and Vietnamese-Lao translation directions, with human post-editing scores used for official rankings.
Discussion (0). Continue with ORCID to comment.