Pith. sign in

REVIEW 1 cited by

Open foundation models for Azerbaijani language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.02337 v2 pith:5XR7DY5P submitted 2024-07-02 cs.CL

classification cs.CL
keywords modelsazerbaijanilanguagefoundationlargeopenopen-sourceseveral
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The emergence of multilingual large language models has enabled the development of language understanding and generation systems in Azerbaijani. However, most of the production-grade systems rely on cloud solutions, such as GPT-4. While there have been several attempts to develop open foundation models for Azerbaijani, these works have not found their way into common use due to a lack of systemic benchmarking. This paper encompasses several lines of work that promote open-source foundation models for Azerbaijani. We introduce (1) a large text corpus for Azerbaijani, (2) a family of encoder-only language models trained on this dataset, (3) labeled datasets for evaluating these models, and (4) extensive evaluation that covers all major open-source models with Azerbaijani support.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TyDi QA-WANA: A Benchmark for Information-Seeking Question Answering in Languages of West Asia and North Africa

    cs.CL 2025-07 conditional novelty 7.0 of 10

    TyDi QA-WANA is a new 28,000-example QA benchmark covering 10 under-represented languages with long-context, information-seeking questions and baseline evaluations.

Pith tools