Pith. sign in

REVIEW 3 major objections 6 minor 8 references

MORFES shows that productive Modern Greek inflection can be measured separately from memorized forms, and that targeted post-training lifts production far above recognition-only competence.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 12:33 UTC pith:VXBRGPIY

load-bearing objection Solid, release-ready Greek morphology benchmark with careful construct design; the Zipf “rule over recall” claim is a proxy, not a proof, and their model lead is real under a fixed protocol but not fully reproducible. the 3 major comments →

arxiv 2607.28274 v1 pith:VXBRGPIY submitted 2026-07-30 cs.CL cs.LG

MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek

classification cs.CL cs.LG
keywords Modern Greekmorphological inflectionproductive competencelanguage model evaluationrecognition vs productionfrequency controlopen-weight modelsbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Modern Greek packs many systematic word forms into nouns, adjectives, and verbs, yet language models for Greek are usually judged on facts, not on whether they can form those words by rule. This paper introduces MORFES, a 500-item expert-verified suite that asks models both to pick and to write the right inflected forms, using lower-frequency lemmas so a correct answer is more likely to mean rule use than rote recall. Recognition is easier than production for every model tested; writing every requested form correctly separates systems much more sharply. The authors’ open model Sophea-Genesis-1 leads on the morphology task while staying comparable to similar-sized open models on English and Greek knowledge, reading, and instruction-following. The release gives the field a dedicated probe for grammatical competence in a morphologically rich language as open models scale multilingual coverage.

Core claim

MORFES is the first expert-curated benchmark dedicated to productive open-class inflection in Modern Greek. With 500 items pairing recognition and production, near-minimal distractors, systematic coverage of inflectional axes, and a preference for lower-frequency lemmas, it scores production per item so that success reflects rule-based formation rather than retrieval of memorized high-frequency forms. Under a fixed 0-shot protocol, Sophea-Genesis-1 reaches 84% per-item production accuracy—well above strong open baselines—while matching similar-scale models on a general-capability panel.

What carries the argument

MORFES itself: a dual-mode (recognition vs production), frequency-controlled, expert-verified item suite over nouns, adjectives, and verbs, with accepted-answer sets that credit any attested correct form but require exact accentuation, and with lemmas released for decontamination.

Load-bearing premise

Preferring lower-frequency lemmas on a subtitle-corpus frequency scale is enough to treat a correct answer as evidence of productive rules rather than memorization from pretraining.

What would settle it

If models that never saw the MORFES lemmas still score high on production only for forms that appear often in large web crawls, or if holding out lemmas fails to drop production while recognition stays high, the claim that the benchmark isolates rule over recall would fail.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Production accuracy per item becomes a sharper yardstick than multiple-choice recognition for Greek morphology in open models.
  • Lemma lists released with the benchmark enable training-time hold-out so future scores can be less contaminated.
  • Targeted post-training can raise inflectional production without trading away general English and Greek capability at similar scale.
  • As open multilingual models grow, MORFES supplies a concrete check on whether scaling yields genuine grammatical competence or only broader memorized coverage.
  • Purpose-built Greek models that recognize well but produce poorly are flagged as likely relying on stored forms rather than internalized paradigms.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The large recognition–production gap across baselines suggests evaluation suites for other fusional languages should weight free-form production, not only choice among candidates.
  • If frequency banding still leaves residual pretraining exposure, pairing MORFES-style items with true nonce lemmas could further isolate rule learning.
  • The same construct—score production on lower-frequency real lemmas with expert accepted sets—could be ported to other morphologically rich lower-resource languages that currently lack dedicated inflection benchmarks.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces MORFES, a 500-item expert-verified benchmark for productive open-class inflection in Modern Greek, pairing recognition (4-way MC) with production (free generation) and preferring lower-frequency lemmas so that correct answers are more plausibly attributed to rule application than memorization. Items cover nouns, adjectives, and verbs (verb-weighted), with near-minimal distractors, multi-form accepted answers, and exact accentuation required. The authors evaluate open models under a fixed 0-shot lm-evaluation-harness/vLLM protocol and report that their post-trained Sophea-Genesis-1 (from Qwen3.6-27B, lemmas held out of morphology post-training) leads on per-item production (84.0%) while remaining comparable on a general-capability panel (MMLU, GreekMMLU, Belebele, IFEval). Dataset and model weights are released publicly.

Significance. If the resource and protocol hold up, MORFES fills a clear gap: Modern Greek LM evaluation has centered on factual knowledge (e.g., GreekMMLU), while existing morphological resources either lack expert curation (SIGMORPHON-style automatic sampling), omit Greek (IMPACT), or test only agreement (MultiBLiMP). A frequency-controlled, production-scored open-class suite is a useful addition for morphologically rich, lower-resource languages, and the public release of both benchmark and Sophea-Genesis-1 weights supports reproducibility and decontamination-by-lemma-holdout. The recognition–production gap is itself a useful diagnostic. Strengths include explicit construct tying (Sections 3.1–3.6), expert verification of accepted-answer sets, and a transparent fixed evaluation harness.

major comments (3)
  1. [Section 3.5] Section 3.5 (and Limitations): The central construct claim—that a correct answer reflects productive rule application rather than recall—rests on SUBTLEX-GR Zipf banding (low ≤3 preferred; verbs via within-voice aggregate of basic finite forms). This is a reasonable proxy but does not establish absence of target cells or related paradigm mates from web-scale pretraining of the baselines. Lemma hold-out is stated only for the authors’ morphology post-training (Section 4.1). For non-author models, ‘rule over recall’ remains an assumption. The paper should either (a) strengthen the claim language to ‘frequency-controlled proxy for reduced memorization risk’ throughout Abstract/Introduction/Conclusion, or (b) add a limited contamination/probe analysis (e.g., surface-form web presence or membership inference on a subset) so the construct validity claim is proportionate to the evidence.
  2. [Section 5] Section 5 / Table 2: The Discussion interprets poor production by Meltemi (13.6%) and Krikri (30.2%) as evidence that ‘current Greek-specialized models may rely more on memorized forms than on internalized rules.’ Those models are 7–8B versus 22–32B for the general open baselines and Sophea-Genesis-1. Size and base-model family are confounded with ‘Greek-specialized training,’ so the memorization diagnosis is not supported by the present design. Either restrict that claim to same-scale comparisons, add a size-matched control discussion, or reframe as ‘under the evaluated scale and data regimes, production lags recognition.’
  3. [Section 4.1] Section 4.1–4.2 / Table 2: Sophea-Genesis-1’s large gain over its base Qwen3.6-27B (84.0% vs 63.0% per-item production) is a headline result, but the post-training data and recipe are proprietary. Without even a high-level description of objective, data mixture type (synthetic paradigms vs. natural text), or scale of morphology-targeted data—while asserting lemma hold-out—the scientific attribution of the gain remains opaque. A brief, non-revealing methods paragraph (data class, hold-out procedure, training stage) would make the leaderboard claim interpretable rather than purely product-announcement.
minor comments (6)
  1. [Section 6] Limitations already notes single-linguist verification and no IAA. Given that accepted forms are largely system-determined, this is acceptable, but a short statement on how borderline multi-form cells (e.g., formal vs. everyday genitives) were adjudicated would help.
  2. [Section 3.5] Table 1: seven nouns/adjectives and seven verbs are n/a on frequency. Briefly state how these were treated relative to the ‘prefer lower-frequency’ policy (included only if unattested = rare, or other rule).
  3. [Appendix A] Appendix A inventories classes with item counts; several rare classes have n=1. The Limitations note that rarest classes are too thin for per-class scores—consider stating explicitly that leaderboard claims are aggregate-only.
  4. [Section 2] Related Work cites MultiBLiMP as Jumelet et al., 2026 and Gemma-4 / GreekMMLU with 2026 dates; ensure bibliography consistency with public versions at camera-ready.
  5. [Section 4.1] Production system prompt is Greek-only and forbids subject pronouns/labels—good for format control. A one-line note on whether any model systematically violated format (and how non-conforming outputs were scored) would improve reproducibility of the per-form metric.
  6. [Abstract] Minor wording: Abstract/Introduction say ‘no benchmark is dedicated to their inflectional competence’—true for expert-curated productive open-class suites, but SIGMORPHON did include Greek automatically; the contrast in Section 2.1 is clearer than the absolute phrasing in the Abstract.

Circularity Check

0 steps flagged

No derivation-by-construction circularity; ordinary author-benchmark self-evaluation only.

full rationale

MORFES is a resource-and-evaluation paper, not a first-principles derivation. The load-bearing claims are (i) that the 500-item suite operationalizes productive open-class inflection via expert curation, near-minimal distractors, axis coverage, and SUBTLEX-GR Zipf preference for lower-frequency lemmas, and (ii) that under a fixed 0-shot lm-evaluation-harness protocol Sophea-Genesis-1 leads per-item production (84.0%) while remaining comparable on the external general-capability panel (MMLU, GreekMMLU, Belebele, IFEval). Neither claim reduces to its inputs by definition: recognition/production accuracies are external behavioral scores against linguist-authored accepted-answer sets, not fitted targets renamed as predictions. The authors explicitly hold the full lemma set out of morphological post-training for Sophea-Genesis-1 (Section 4.1), which breaks train-on-test circularity for their model. Frequency banding is a construct-validity assumption about pretraining memorization, not a circular step. Residual risk is only the ordinary fact that the same lab releases both the benchmark and the leading model—an evaluation-integrity concern, not a self-definitional or fitted-input loop. No self-citation uniqueness theorem, smuggled ansatz, or renamed known result carries the central claim. Score 1 reflects that mild self-evaluation proximity only; steps are empty because no enumerated circular reduction is present.

Axiom & Free-Parameter Ledger

3 free parameters · 6 axioms · 2 invented entities

As an empirical benchmark paper, load-bearing commitments are construct/operationalization choices and corpus frequency assumptions, not fitted physical constants. The central leaderboard claim rests on accepting those measurement axioms and the fixed eval protocol.

free parameters (3)
  • Zipf low/transitional/high cutoffs (≤3, 3–4, ≥4) = low≤3, transitional 3–4, high≥4
    Hand-chosen frequency bands that define which lemmas count as ‘lower-frequency’ and thus support the rule-not-recall interpretation.
  • Item mix (300 verbs / 100 nouns / 100 adjectives; half single-form half full-paradigm for verbs) = 300/100/100; 150+150 verb split
    Design allocation that weights the overall score toward verbal morphology; different mix would change aggregate rankings.
  • Verb frequency aggregate (sum of selected basic finite forms within one voice) = present+imperfect+aorist within-voice sum → Zipf
    Author-chosen proxy replacing dictionary-form frequency for verbs; changes which verbs land in low vs transitional bands.
axioms (6)
  • domain assumption Productive inflectional competence is validly measured by exact-form recognition and production on real lower-frequency open-class lemmas with near-minimal distractors.
    Section 3.1 ties every design decision to this construct; alternative constructs (nonce words, agreement only, GEC) are rejected.
  • domain assumption SUBTLEX-GR subtitle frequencies are an adequate proxy for likelihood that a form was memorized in LM pretraining.
    Section 3.5 bases frequency control on this corpus and Zipf scaling; pretraining corpora differ in domain and scale.
  • domain assumption Closed-class and adverb inflection should be excluded because they can be memorized or are too thin to evidence rich productive paradigms.
    Section 3.1 scope restriction shapes what ‘inflectional competence’ means on MORFES.
  • domain assumption Any linguist-accepted variant (including optional article) with exact accentuation is a correct production; stress errors are always wrong.
    Section 3.6 scoring rule directly defines production accuracy.
  • ad hoc to paper Length-normalized log-likelihood ranking of raw completions without chat template is a fair recognition metric across models.
    Section 4.1 evaluation protocol; different likelihood normalizations or chat-template use could reorder recognition scores.
  • domain assumption Standard categorical grammar inventory from Chatzisavvidis & Chatzisavvidou (2011) is a sufficient checklist of classes to cover.
    Section 3.4 and Appendix A; coverage claims depend on that inventory.
invented entities (2)
  • MORFES construct operationalization (paired recognition/production suite with frequency-controlled open-class items) independent evidence
    purpose: Turn ‘productive inflectional competence’ into a 500-item scored benchmark.
    Not a physical entity but a new measurement object; independent use is possible via public HF dataset, so it is falsifiable by external re-annotation or contamination studies.
  • Sophea-Genesis-1 (post-trained Qwen3.6-27B open weights) independent evidence
    purpose: Demonstrate that targeted post-training can raise MORFES production while preserving general panel scores.
    New model artifact released on HF; training mixture proprietary, so external groups can evaluate but not fully retrain identically.

pith-pipeline@v1.2.0-daily-grok45 · 17916 in / 3895 out tokens · 81549 ms · 2026-07-31T12:33:29.497596+00:00 · methodology

0 comments
read the original abstract

Modern Greek is a richly inflected language, yet the language models built for it are evaluated mainly on factual knowledge, and no benchmark is dedicated to their inflectional competence. We introduce MORFES (Morphological Open-class Recognition-and-Formation Evaluation Suite), a benchmark of 500 expert-verified items that tests the recognition and production of Greek inflected forms, favoring lower-frequency lemmas so that a correct answer reflects the rule rather than a memorized form. We make it publicly available at https://huggingface.co/datasets/KIEFERSA/MORFES. We evaluate a range of open language models on MORFES, situating them within the rapidly scaling open-weight ecosystem from LLaMA to Qwen3, DeepSeek-R1, Magistral, and Kimi K2, where multilingual coverage grows but grammatical competence in morphologically rich languages remains under-measured. Among them, Sophea-Genesis-1, a model we developed and release as open weights at https://huggingface.co/KIEFERSA/Sophea-Genesis-1, leads on inflectional morphology while matching similarly sized models in general capability.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

8 extracted references · 6 linked inside Pith

  1. [2]

    In Advances in Neural Information Processing Systems (NeurIPS 2025), Datasets and Benchmarks Track

    Measuring what matters: Construct validity in large lan- guage model benchmarks. In Advances in Neural Information Processing Systems (NeurIPS 2025), Datasets and Benchmarks Track. Sofronis Chatzisavvidis and Athanasia Chatzisavvidou. 2011. Γραμματική Νέας Ελληνικής Γλώσσας [Grammar of Modern Greek]. Organization for the Publication of Educational Books (...

  2. [5]

    arXiv preprint arXiv:2505.13772

    Krikri: Advancing open large language models for Greek. arXiv preprint arXiv:2505.13772. Mohammed J. Saeed, Tommi Vehvilainen, Evgeny Fedoseev, Sevil Caliskan, and Tatiana Vodolazova. 2025. IMPACT: Inflectional morphology probes across complex typologies. arXiv preprint arXiv:2506.23929. 8 Robert Schreuder and R. Harald Baayen. 1995. Modeling morpho- logi...

  3. [6]

    arXiv preprint arXiv:2407.20743

    Meltemi: The first open large language model for Greek. arXiv preprint arXiv:2407.20743. Leonie Weissweiler, Valentin Hofmann, Anjali Kantharuban, et al

  4. [8]

    I write”); in the passive voice, marked by the ending ­μαι, the subject is instead on the receiving end of the action, and can also carry a reflexive (“to oneself

    GreekMMLU: A native-sourced multitask benchmark for evaluating language models in Greek. In Findings of the Association for Computational Linguistics: ACL 2026. Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, et al. 2023. Instruc - tion-following evaluation for large language models. arXiv preprint arXiv:2311.07911. 9 Appendix A: Inflectional-Class Inventory T...

  5. [2023]

    In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6508–6524, Singapore

    Counting the bugs in ChatGPT’s wugs: A multilingual investigation into the morphological capabilities of a large language model. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6508–6524, Singapore. Yang Zhang, Mersin Konomi, Christos Xypolopoulos, et al

  6. [2024]

    Gemma Team

    The language model evaluation harness. Gemma Team. 2026. Gemma 4 technical report. arXiv preprint arXiv:2607.02770. Omer Goldman, David Guriel, and Reut Tsarfaty. 2022. (Un)solv- ing morphological inflection: Lemma overlap artificially inflates models’ performance. In Proceedings of the 60th Annual Meet­ ing of the Association for Computational Linguistic...

  7. [2025]

    In Tokenization Workshop at ICML 2025

    Evaluating morphological alignment of tokenizers in 70 languages. In Tokenization Workshop at ICML 2025 . arXiv:2507.06378. Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinan- dan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2024. The Belebele benchmark: A parallel reading comprehension ...

  8. [2026]

    arXiv preprint arXiv:2602.05879

    EuroLLM-22B: Technical report. arXiv preprint arXiv:2602.05879. Dimitris Roussis, Leon Voukoutis, Georgios Paraskevopoulos, et al