A breadth-first survey of recent South Asian language NLP, speech, and multimodal research, organized with LLM-based classification and clustering.
Performance Evaluation of Tokenizers in Large Language Models for the Assamese Language
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Training of a tokenizer plays an important role in the performance of deep learning models. This research aims to understand the performance of tokenizers in five state-of-the-art (SOTA) large language models (LLMs) in the Assamese language of India. The research is important to understand the multi-lingual support for a low-resourced language such as Assamese. Our research reveals that the tokenizer of SUTRA from Two AI performs the best with an average Normalized Sequence Length (NSL) value of 0.45, closely followed by the tokenizer of GPT-4o from Open AI with an average NSL value of 0.54, followed by Gemma 2, Meta Llama 3.1, and Mistral Large Instruct 2407 with an average NSL value of 0.82, 1.4, and 1.48 respectively.
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Breadth-First Catalog of Text Processing, Speech Processing and Multimodal Research in South Asian Languages
A breadth-first survey of recent South Asian language NLP, speech, and multimodal research, organized with LLM-based classification and clustering.