Pith. sign in

REVIEW 2 cited by

Performance Evaluation of Tokenizers in Large Language Models for the Assamese Language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.03718 v1 pith:K4DECKON submitted 2024-09-28 cs.CL

classification cs.CL
keywords languageassameseaveragelargemodelsperformanceresearchtokenizer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training of a tokenizer plays an important role in the performance of deep learning models. This research aims to understand the performance of tokenizers in five state-of-the-art (SOTA) large language models (LLMs) in the Assamese language of India. The research is important to understand the multi-lingual support for a low-resourced language such as Assamese. Our research reveals that the tokenizer of SUTRA from Two AI performs the best with an average Normalized Sequence Length (NSL) value of 0.45, closely followed by the tokenizer of GPT-4o from Open AI with an average NSL value of 0.54, followed by Gemma 2, Meta Llama 3.1, and Mistral Large Instruct 2407 with an average NSL value of 0.82, 1.4, and 1.48 respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Breadth-First Catalog of Text Processing, Speech Processing and Multimodal Research in South Asian Languages

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A breadth-first survey of recent South Asian language NLP, speech, and multimodal research, organized with LLM-based classification and clustering.

  2. Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages

    cs.CL 2024-11 reject novelty 4.0 of 10

    A single-sentence evaluation claims SUTRA's tokenizer is most token-efficient across 14 of India's 22 official languages, based on one example text per language.

Pith tools