Benchmarking Llama-70b against GPT-4o mini on Greek shows task-specific strengths, but the contamination-probe and legal-clustering claims need stronger baselines.
A Systematic Survey of Natural Language Processing for the Greek Language
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Comprehensive monolingual Natural Language Processing (NLP) surveys are essential for assessing language-specific challenges, resource availability, and research gaps. However, existing surveys often lack standardized methodologies, leading to selection bias and fragmented coverage of NLP tasks and resources. This study introduces a generalizable framework for systematic monolingual NLP surveys. Our approach integrates a structured search protocol to minimize bias, an NLP task taxonomy for classification, and language resource taxonomies to identify potential benchmarks and highlight opportunities for improving resource availability. We apply this framework to Greek NLP (2012-2023), providing an in-depth analysis of its current state, task-specific progress, and resource gaps. The survey results are publicly available (https://doi.org/10.5281/zenodo.15314882) and are regularly updated to provide an evergreen resource. This systematic survey of Greek NLP serves as a case study, demonstrating the effectiveness of our framework and its potential for broader application to other not so well-resourced languages as regards NLP.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Open or Closed LLM for Lesser-Resourced Languages? Lessons from Greek
Benchmarking Llama-70b against GPT-4o mini on Greek shows task-specific strengths, but the contamination-probe and legal-clustering claims need stronger baselines.