Pith. sign in

REVIEW 4 cited by

Meltemi: The first open Large Language Model for Greek

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.20743 v1 pith:HE43VQPA submitted 2024-07-30 cs.CL

classification cs.CL
keywords meltemigreekcorpusinstructlanguagemodelbeenbillion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We describe the development and capabilities of Meltemi 7B, the first open Large Language Model for the Greek language. Meltemi 7B has 7 billion parameters and is trained on a 40 billion token Greek corpus. For the development of Meltemi 7B, we adapt Mistral, by continuous pretraining on the Greek Corpus. Meltemi 7B contains up-to-date information up to September 2023. Furthermore, we have translated and curated a Greek instruction corpus, which has been used for the instruction-tuning of a chat model, named Meltemi 7B Instruct. Special care has been given to the alignment and the removal of toxic content for the Meltemi 7B Instruct. The developed models are evaluated on a broad set of collected evaluation corpora, and examples of prompts and responses are presented. Both Meltemi 7B and Meltemi 7B Instruct are available at https://huggingface.co/ilsp under the Apache 2.0 license.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek

    cs.CL 2026-07 conditional novelty 6.0 of 10

    MORFES is the first expert-verified Modern Greek productive-inflection benchmark; Sophea-Genesis-1 leads it at 84% per-item production without losing general capability.

  2. Open or Closed LLM for Lesser-Resourced Languages? Lessons from Greek

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Benchmarking Llama-70b against GPT-4o mini on Greek shows task-specific strengths, but the contamination-probe and legal-clustering claims need stronger baselines.

  3. SnakModel: Lessons Learned from Training an Open Danish Large Language Model

    cs.CL 2024-12 conditional novelty 5.0 of 10

    SnakModel, an open Danish 7B LLM, outperforms other Llama2-7B-based models on the ScandEval Danish benchmark, with analyses of training dynamics and data curation.

  4. GR-NLP-TOOLKIT: An Open-Source NLP Toolkit for Modern Greek

    cs.CL 2024-12 conditional novelty 4.0 of 10

    GR-NLP-TOOLKIT is a pip-installable Greek NLP toolkit reporting improved scores over spaCy and Stanza on five tasks, built by fine-tuning GreekBERT and ByT5.

Pith tools