Pith. sign in

REVIEW 1 cited by

Tamil-Llama: A New Tamil Language Model Based on Llama 2

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.05845 v1 pith:FZV5UWAE submitted 2023-11-10 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords tamillanguagemodelgenerationmodelstextdatasetfurther
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language modeling has witnessed remarkable advancements in recent years, with Large Language Models (LLMs) like ChatGPT setting unparalleled benchmarks in human-like text generation. However, a prevailing limitation is the underrepresentation of languages like Tamil in these cutting-edge models, leading to suboptimal performance in diverse linguistic contexts. This paper addresses this lacuna, enhancing the open-source LLaMA model with an addition of 16,000 Tamil tokens, aiming to achieve superior text generation and comprehension in the Tamil language. We strategically employ the LoRA methodology for efficient model training on a comprehensive Tamil corpus, ensuring computational feasibility and model robustness. Moreover, we introduce a Tamil-translated version of the Alpaca dataset and a subset of the OpenOrca dataset tailored for instruction fine-tuning. Our results showcase significant performance improvements in Tamil text generation, with potential implications for the broader landscape of LLMs in Indian languages. We further underscore our commitment to open research by making our models, datasets, and code publicly accessible, fostering further innovations in language modeling.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. It's All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMs

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Zero-shot in-context learning with word- or sentence-level English alignment outperforms fine-tuning for LLMs on extremely low-resource languages with rare scripts.

Pith tools