Pith. sign in

REVIEW 4 cited by

VinaLLaMA: LLaMA-based Vietnamese Foundation Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.11011 v1 pith:PCFABEED submitted 2023-12-18 cs.CL

classification cs.CL
keywords vietnamesemodelvinallamalanguagesotatrainedachievesadditional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this technical report, we present VinaLLaMA, an open-weight, state-of-the-art (SOTA) Large Language Model for the Vietnamese language, built upon LLaMA-2 with an additional 800 billion trained tokens. VinaLLaMA not only demonstrates fluency in Vietnamese but also exhibits a profound understanding of Vietnamese culture, making it a truly indigenous model. VinaLLaMA-7B-chat, trained on 1 million high-quality synthetic samples, achieves SOTA results on key benchmarks, including VLSP, VMLU, and Vicuna Benchmark Vietnamese, marking a significant advancement in the Vietnamese AI landscape and offering a versatile resource for various applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Soro: A Lightweight Foundation Model and Chatbot for Tajik

    cs.AI 2026-04 conditional novelty 5.5 of 10

    Tajik-specialized Gemma 3 derivatives (12B/27B) beat same-size baselines by ~6–8 points on new Tajik exams after 1.9B-token continual pretraining, with FP8/INT4 still usable on edge GPUs.

  2. VietMed-MCQ: A Consistency-Filtered Data Synthesis Framework for Vietnamese Traditional Medicine Evaluation

    cs.CL 2026-01 conditional novelty 5.0 of 10

    A dual-model consistency filter and a 94.2% expert-approved sample back VietMed-MCQ, a 3,190-question Vietnamese Traditional Medicine quiz where Qwen2.5 models beat Vietnamese-specialized models by 7.2%.

  3. ViBidirectionMT-Eval: Machine Translation for Vietnamese-Chinese and Vietnamese-Lao language pair

    cs.CL 2025-01 reject novelty 4.0 of 10

    A shared-task report and dataset release for four Vietnamese-Chinese and Vietnamese-Lao translation directions, with human post-editing scores used for official rankings.

  4. BgGPT 1.0: Extending English-centric LLMs to other languages

    cs.CL 2024-12 conditional novelty 4.0 of 10

    Continually pretraining Gemma-2 on a curated Bulgarian corpus and merging with instruction-tuned models yields open Bulgarian-English models that beat larger open models on Bulgarian benchmarks.

Pith tools