REVIEW 3 major objections 6 minor 8 references
MORFES shows that productive Modern Greek inflection can be measured separately from memorized forms, and that targeted post-training lifts production far above recognition-only competence.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 12:33 UTC pith:VXBRGPIY
load-bearing objection Solid, release-ready Greek morphology benchmark with careful construct design; the Zipf “rule over recall” claim is a proxy, not a proof, and their model lead is real under a fixed protocol but not fully reproducible. the 3 major comments →
MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
MORFES is the first expert-curated benchmark dedicated to productive open-class inflection in Modern Greek. With 500 items pairing recognition and production, near-minimal distractors, systematic coverage of inflectional axes, and a preference for lower-frequency lemmas, it scores production per item so that success reflects rule-based formation rather than retrieval of memorized high-frequency forms. Under a fixed 0-shot protocol, Sophea-Genesis-1 reaches 84% per-item production accuracy—well above strong open baselines—while matching similar-scale models on a general-capability panel.
What carries the argument
MORFES itself: a dual-mode (recognition vs production), frequency-controlled, expert-verified item suite over nouns, adjectives, and verbs, with accepted-answer sets that credit any attested correct form but require exact accentuation, and with lemmas released for decontamination.
Load-bearing premise
Preferring lower-frequency lemmas on a subtitle-corpus frequency scale is enough to treat a correct answer as evidence of productive rules rather than memorization from pretraining.
What would settle it
If models that never saw the MORFES lemmas still score high on production only for forms that appear often in large web crawls, or if holding out lemmas fails to drop production while recognition stays high, the claim that the benchmark isolates rule over recall would fail.
If this is right
- Production accuracy per item becomes a sharper yardstick than multiple-choice recognition for Greek morphology in open models.
- Lemma lists released with the benchmark enable training-time hold-out so future scores can be less contaminated.
- Targeted post-training can raise inflectional production without trading away general English and Greek capability at similar scale.
- As open multilingual models grow, MORFES supplies a concrete check on whether scaling yields genuine grammatical competence or only broader memorized coverage.
- Purpose-built Greek models that recognize well but produce poorly are flagged as likely relying on stored forms rather than internalized paradigms.
Where Pith is reading between the lines
- The large recognition–production gap across baselines suggests evaluation suites for other fusional languages should weight free-form production, not only choice among candidates.
- If frequency banding still leaves residual pretraining exposure, pairing MORFES-style items with true nonce lemmas could further isolate rule learning.
- The same construct—score production on lower-frequency real lemmas with expert accepted sets—could be ported to other morphologically rich lower-resource languages that currently lack dedicated inflection benchmarks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MORFES, a 500-item expert-verified benchmark for productive open-class inflection in Modern Greek, pairing recognition (4-way MC) with production (free generation) and preferring lower-frequency lemmas so that correct answers are more plausibly attributed to rule application than memorization. Items cover nouns, adjectives, and verbs (verb-weighted), with near-minimal distractors, multi-form accepted answers, and exact accentuation required. The authors evaluate open models under a fixed 0-shot lm-evaluation-harness/vLLM protocol and report that their post-trained Sophea-Genesis-1 (from Qwen3.6-27B, lemmas held out of morphology post-training) leads on per-item production (84.0%) while remaining comparable on a general-capability panel (MMLU, GreekMMLU, Belebele, IFEval). Dataset and model weights are released publicly.
Significance. If the resource and protocol hold up, MORFES fills a clear gap: Modern Greek LM evaluation has centered on factual knowledge (e.g., GreekMMLU), while existing morphological resources either lack expert curation (SIGMORPHON-style automatic sampling), omit Greek (IMPACT), or test only agreement (MultiBLiMP). A frequency-controlled, production-scored open-class suite is a useful addition for morphologically rich, lower-resource languages, and the public release of both benchmark and Sophea-Genesis-1 weights supports reproducibility and decontamination-by-lemma-holdout. The recognition–production gap is itself a useful diagnostic. Strengths include explicit construct tying (Sections 3.1–3.6), expert verification of accepted-answer sets, and a transparent fixed evaluation harness.
major comments (3)
- [Section 3.5] Section 3.5 (and Limitations): The central construct claim—that a correct answer reflects productive rule application rather than recall—rests on SUBTLEX-GR Zipf banding (low ≤3 preferred; verbs via within-voice aggregate of basic finite forms). This is a reasonable proxy but does not establish absence of target cells or related paradigm mates from web-scale pretraining of the baselines. Lemma hold-out is stated only for the authors’ morphology post-training (Section 4.1). For non-author models, ‘rule over recall’ remains an assumption. The paper should either (a) strengthen the claim language to ‘frequency-controlled proxy for reduced memorization risk’ throughout Abstract/Introduction/Conclusion, or (b) add a limited contamination/probe analysis (e.g., surface-form web presence or membership inference on a subset) so the construct validity claim is proportionate to the evidence.
- [Section 5] Section 5 / Table 2: The Discussion interprets poor production by Meltemi (13.6%) and Krikri (30.2%) as evidence that ‘current Greek-specialized models may rely more on memorized forms than on internalized rules.’ Those models are 7–8B versus 22–32B for the general open baselines and Sophea-Genesis-1. Size and base-model family are confounded with ‘Greek-specialized training,’ so the memorization diagnosis is not supported by the present design. Either restrict that claim to same-scale comparisons, add a size-matched control discussion, or reframe as ‘under the evaluated scale and data regimes, production lags recognition.’
- [Section 4.1] Section 4.1–4.2 / Table 2: Sophea-Genesis-1’s large gain over its base Qwen3.6-27B (84.0% vs 63.0% per-item production) is a headline result, but the post-training data and recipe are proprietary. Without even a high-level description of objective, data mixture type (synthetic paradigms vs. natural text), or scale of morphology-targeted data—while asserting lemma hold-out—the scientific attribution of the gain remains opaque. A brief, non-revealing methods paragraph (data class, hold-out procedure, training stage) would make the leaderboard claim interpretable rather than purely product-announcement.
minor comments (6)
- [Section 6] Limitations already notes single-linguist verification and no IAA. Given that accepted forms are largely system-determined, this is acceptable, but a short statement on how borderline multi-form cells (e.g., formal vs. everyday genitives) were adjudicated would help.
- [Section 3.5] Table 1: seven nouns/adjectives and seven verbs are n/a on frequency. Briefly state how these were treated relative to the ‘prefer lower-frequency’ policy (included only if unattested = rare, or other rule).
- [Appendix A] Appendix A inventories classes with item counts; several rare classes have n=1. The Limitations note that rarest classes are too thin for per-class scores—consider stating explicitly that leaderboard claims are aggregate-only.
- [Section 2] Related Work cites MultiBLiMP as Jumelet et al., 2026 and Gemma-4 / GreekMMLU with 2026 dates; ensure bibliography consistency with public versions at camera-ready.
- [Section 4.1] Production system prompt is Greek-only and forbids subject pronouns/labels—good for format control. A one-line note on whether any model systematically violated format (and how non-conforming outputs were scored) would improve reproducibility of the per-form metric.
- [Abstract] Minor wording: Abstract/Introduction say ‘no benchmark is dedicated to their inflectional competence’—true for expert-curated productive open-class suites, but SIGMORPHON did include Greek automatically; the contrast in Section 2.1 is clearer than the absolute phrasing in the Abstract.
Circularity Check
No derivation-by-construction circularity; ordinary author-benchmark self-evaluation only.
full rationale
MORFES is a resource-and-evaluation paper, not a first-principles derivation. The load-bearing claims are (i) that the 500-item suite operationalizes productive open-class inflection via expert curation, near-minimal distractors, axis coverage, and SUBTLEX-GR Zipf preference for lower-frequency lemmas, and (ii) that under a fixed 0-shot lm-evaluation-harness protocol Sophea-Genesis-1 leads per-item production (84.0%) while remaining comparable on the external general-capability panel (MMLU, GreekMMLU, Belebele, IFEval). Neither claim reduces to its inputs by definition: recognition/production accuracies are external behavioral scores against linguist-authored accepted-answer sets, not fitted targets renamed as predictions. The authors explicitly hold the full lemma set out of morphological post-training for Sophea-Genesis-1 (Section 4.1), which breaks train-on-test circularity for their model. Frequency banding is a construct-validity assumption about pretraining memorization, not a circular step. Residual risk is only the ordinary fact that the same lab releases both the benchmark and the leading model—an evaluation-integrity concern, not a self-definitional or fitted-input loop. No self-citation uniqueness theorem, smuggled ansatz, or renamed known result carries the central claim. Score 1 reflects that mild self-evaluation proximity only; steps are empty because no enumerated circular reduction is present.
Axiom & Free-Parameter Ledger
free parameters (3)
- Zipf low/transitional/high cutoffs (≤3, 3–4, ≥4) =
low≤3, transitional 3–4, high≥4
- Item mix (300 verbs / 100 nouns / 100 adjectives; half single-form half full-paradigm for verbs) =
300/100/100; 150+150 verb split
- Verb frequency aggregate (sum of selected basic finite forms within one voice) =
present+imperfect+aorist within-voice sum → Zipf
axioms (6)
- domain assumption Productive inflectional competence is validly measured by exact-form recognition and production on real lower-frequency open-class lemmas with near-minimal distractors.
- domain assumption SUBTLEX-GR subtitle frequencies are an adequate proxy for likelihood that a form was memorized in LM pretraining.
- domain assumption Closed-class and adverb inflection should be excluded because they can be memorized or are too thin to evidence rich productive paradigms.
- domain assumption Any linguist-accepted variant (including optional article) with exact accentuation is a correct production; stress errors are always wrong.
- ad hoc to paper Length-normalized log-likelihood ranking of raw completions without chat template is a fair recognition metric across models.
- domain assumption Standard categorical grammar inventory from Chatzisavvidis & Chatzisavvidou (2011) is a sufficient checklist of classes to cover.
invented entities (2)
-
MORFES construct operationalization (paired recognition/production suite with frequency-controlled open-class items)
independent evidence
-
Sophea-Genesis-1 (post-trained Qwen3.6-27B open weights)
independent evidence
read the original abstract
Modern Greek is a richly inflected language, yet the language models built for it are evaluated mainly on factual knowledge, and no benchmark is dedicated to their inflectional competence. We introduce MORFES (Morphological Open-class Recognition-and-Formation Evaluation Suite), a benchmark of 500 expert-verified items that tests the recognition and production of Greek inflected forms, favoring lower-frequency lemmas so that a correct answer reflects the rule rather than a memorized form. We make it publicly available at https://huggingface.co/datasets/KIEFERSA/MORFES. We evaluate a range of open language models on MORFES, situating them within the rapidly scaling open-weight ecosystem from LLaMA to Qwen3, DeepSeek-R1, Magistral, and Kimi K2, where multilingual coverage grows but grammatical competence in morphologically rich languages remains under-measured. Among them, Sophea-Genesis-1, a model we developed and release as open weights at https://huggingface.co/KIEFERSA/Sophea-Genesis-1, leads on inflectional morphology while matching similarly sized models in general capability.
Reference graph
Works this paper leans on
-
[2]
In Advances in Neural Information Processing Systems (NeurIPS 2025), Datasets and Benchmarks Track
Measuring what matters: Construct validity in large lan- guage model benchmarks. In Advances in Neural Information Processing Systems (NeurIPS 2025), Datasets and Benchmarks Track. Sofronis Chatzisavvidis and Athanasia Chatzisavvidou. 2011. Γραμματική Νέας Ελληνικής Γλώσσας [Grammar of Modern Greek]. Organization for the Publication of Educational Books (...
Pith/arXiv arXiv 2025
-
[5]
arXiv preprint arXiv:2505.13772
Krikri: Advancing open large language models for Greek. arXiv preprint arXiv:2505.13772. Mohammed J. Saeed, Tommi Vehvilainen, Evgeny Fedoseev, Sevil Caliskan, and Tatiana Vodolazova. 2025. IMPACT: Inflectional morphology probes across complex typologies. arXiv preprint arXiv:2506.23929. 8 Robert Schreuder and R. Harald Baayen. 1995. Modeling morpho- logi...
Pith/arXiv arXiv 2025
-
[6]
arXiv preprint arXiv:2407.20743
Meltemi: The first open large language model for Greek. arXiv preprint arXiv:2407.20743. Leonie Weissweiler, Valentin Hofmann, Anjali Kantharuban, et al
-
[8]
GreekMMLU: A native-sourced multitask benchmark for evaluating language models in Greek. In Findings of the Association for Computational Linguistics: ACL 2026. Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, et al. 2023. Instruc - tion-following evaluation for large language models. arXiv preprint arXiv:2311.07911. 9 Appendix A: Inflectional-Class Inventory T...
Pith/arXiv arXiv 2026
-
[2023]
In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6508–6524, Singapore
Counting the bugs in ChatGPT’s wugs: A multilingual investigation into the morphological capabilities of a large language model. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6508–6524, Singapore. Yang Zhang, Mersin Konomi, Christos Xypolopoulos, et al
2023
-
[2024]
The language model evaluation harness. Gemma Team. 2026. Gemma 4 technical report. arXiv preprint arXiv:2607.02770. Omer Goldman, David Guriel, and Reut Tsarfaty. 2022. (Un)solv- ing morphological inflection: Lemma overlap artificially inflates models’ performance. In Proceedings of the 60th Annual Meet ing of the Association for Computational Linguistic...
Pith/arXiv arXiv 2026
-
[2025]
In Tokenization Workshop at ICML 2025
Evaluating morphological alignment of tokenizers in 70 languages. In Tokenization Workshop at ICML 2025 . arXiv:2507.06378. Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinan- dan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2024. The Belebele benchmark: A parallel reading comprehension ...
Pith/arXiv arXiv 2025
-
[2026]
arXiv preprint arXiv:2602.05879
EuroLLM-22B: Technical report. arXiv preprint arXiv:2602.05879. Dimitris Roussis, Leon Voukoutis, Georgios Paraskevopoulos, et al
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.