Pith. sign in

REVIEW 4 major objections 5 minor 36 references

Annotating Compositionality Scores for Irish Noun Compounds is Hard Work

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper creates the first collection of Irish noun compounds with compositionality ratings and argues it is a reliable basis for Irish-specific noun-compound resources and for evaluating language models in Irish.

desk verdict First Irish noun-compound compositionality annotation effort, honestly reported, but the claimed dataset is not actually accessible from the paper. read the letter →

arxiv 2502.10061 v1 pith:BKQMPEWL submitted 2025-02-14 cs.CL

classification cs.CL
keywords nouncompoundscompositionalityIrishlanguageannotationguidelinesmultiwordexpressionsdomainspecificityinter-annotatoragreementlow-resourceNLP
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is about making the notion of a noun compound measurable for Irish, a language with almost no such annotated resources. The authors define an Irish noun compound narrowly—as a contiguous two-word construction whose head is a noun, with interleaved determiners and named entities excluded or scored as opaque—and they build annotation guidelines plus a pilot corpus from two very different text sources: dialectal folklore transcriptions and a modern dependency treebank. Their central claim is that expert annotators can rate compositionality, domain specificity, familiarity, and confidence well enough to support corpus building, and that the released pilot data constitute the first collection of Irish noun compounds with compositionality ratings. The paper reports that opaque compounds are roughly three times more frequent in the folklore data than in the modern treebank, and that domain-specificity and confidence scores are lower there, suggesting the two sources capture different populations of compounds. The reason to care is that this is the evidentiary base needed to test whether language models handle idiomatic Irish as well as they handle English.

What carries the argument

The operative object is the noun compound candidate (NCC), defined in the guidelines as a contiguous two-word construction in which one component is a noun or an adjective dependent on a noun head, and the whole construction has nominal distribution. Interleaved determiners disqualify a candidate; named entities can be annotated but receive compositionality 0 by default. The measuring instruments are the six-point compositionality scale (0 opaque to 5 transparent) and the three-point domain-specificity, familiarity, and confidence scales. The pilot-task refinement loop—three rounds with pair-wise weighted kappa values mostly between 0.3 and 0.64—is the procedural machinery that turned initial disagreement into a usable guideline document.

What would settle it

The reported folklore-versus-treebank contrast would be settled by re-annotating the same sentences under a wider definition that includes interleaved-determiner constructions such as 'mí na meala' (honeymoon); if the opaque-compound ratio moves substantially toward the folklore figure, the operational definition is producing the observed difference.

Watch

Extended reading notes

Core claim

The paper's central claim is that noun-compound annotation for Irish is tractable and informative once the object is defined narrowly enough and the guidelines are shaped by pilot disagreements. Working from a definition of a noun compound as a contiguous two-word phrase whose head is a noun and whose distribution is nominal, the authors show that expert annotators can assign six-point compositionality scores together with domain-specificity, familiarity, and confidence ratings, and that the resulting judgements reveal a systematic contrast between text types: the dialectal folklore corpus contains a higher share of opaque compounds (average ratio 0.36 versus 0.12 in the modern treebank) and receives lower domain-specificity and confidence scores. They also report that translating English noun compounds into Irish rarely produces a two-word Irish compound, since many English compounds map to single Irish words or to possessive constructions with interleaved articles. The released pilot annotations and the associated guidelines are offered as the first collection of Irish noun compounds with compositionality ratings.

Load-bearing premise

The load-bearing premise is that defining an Irish noun compound as a contiguous two-word phrase headed by a noun, with interleaved determiners and named entities excluded or scored as opaque, does not bias the measured compositionality away from the constructions Irish speakers actually use.

Editorial extensions

If this is right

  • If the pilot corpus and guidelines are reliable, they provide the first Irish-language benchmark for testing whether language models' handling of idiomatic compounds transfers beyond English.
  • The observed contrast between folklore and modern treebank data implies that any Irish noun-compound resource must be domain-balanced, because folklore-style text over-represents opaque compounds while modern text over-represents compositional ones.
  • The translation finding—only 30 of 280 English noun compounds map to two-word Irish compounds—implies that English compound lexicons cannot simply be projected onto Irish; Irish resources have to be built from Irish text.
  • The average inter-annotator agreement around 0.5 weighted kappa implies that compositionality scores should be accompanied by confidence and familiarity metadata, and that downstream uses should treat the scores as graded rather than categorical.
  • The deliberate exclusion of interleaved-determiner constructions narrows the current resource but also marks the boundary for a natural future extension.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the operational definition (contiguous two words, noun head, no interleaved determiner) likely under-counts possessive and article-bearing Irish constructions that carry non-compositional meaning, so the released statistics may understate the proportion of opaque expressions in ordinary Irish text.
  • Editorial inference: a direct testable next step is to feed the released noun compound candidates to multilingual language models and compare their compositionality judgements with the human scores, which would reveal whether Irish opacity is harder for models than the English and Portuguese cases studied elsewhere.
  • Editorial inference: the translation result points to a linguistic asymmetry—English tends to pack non-compositional meanings into noun-plus-noun compounds, while Irish tends to realize them as single words or possessive phrases—which would change how cross-lingual multiword-expression identification should be designed.
  • Editorial inference: the kappa pattern across pilots suggests that annotator training on hard cases matters more than the choice of rating scale, so a small follow-up study varying annotator background and guideline examples could test whether the guidelines transfer beyond the original annotator pool.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reports ongoing annotation work for Irish noun compounds, presenting guidelines for identifying noun compound candidates (NCCs) and for scoring compositionality, domain specificity, annotator familiarity, and confidence. The authors describe three pilot tasks with expert annotators, report Cohen's weighted kappa agreement scores, discuss difficult cases such as definite-article constructions and named entities, and give preliminary statistics contrasting the Dúchas folklore corpus with the UD-IDT treebank. The paper claims to provide the first collection of Irish noun compounds with compositionality ratings and to make pilot annotations publicly available.

Significance. If the dataset and guidelines were released, the work would be a useful first resource for Irish multiword-expression research and for evaluating language models on Irish noun compounds. The use of multiple expert annotators with different dialect backgrounds and the explicit discussion of annotation difficulties are valuable. The paper is honest about its scope limitations, such as restricting NCCs to two-word contiguous constructions and excluding definite-article constructions. However, the central contribution is a resource that is not currently available, and the inter-annotator agreement is moderate, so the empirical claims cannot yet be fully verified.

major comments (4)
  1. [Section 1 and Section 6] The paper states in Section 1 that 'the pilot task annotations are made available for public use', but no URL, DOI, repository, or appendix is provided; Section 6 instead says the collection 'will be released alongside the annotated corpus'. Since the central claim is the creation of a first-of-its-kind Irish noun-compound resource with compositionality ratings, the absence of any accessible artifact makes the core empirical claims unverifiable. The final version must include a working link or DOI and ideally the full item-level annotations.
  2. [Table 1 and Section 4.2] The weighted kappa values in Table 1 range from 0.30 to 0.64, with many pairs below 0.55, which is generally considered only moderate or weak agreement. The text in Section 4.2 says the agreement results 'provide a glimpse into the level of consensus', but the later framing that the guidelines support 'reliable annotation' is not supported by these numbers. Please report the weighting scheme used, confidence intervals, and the number of items per pilot, and discuss whether the pilot tasks were used iteratively to revise the guidelines or as a final reliability benchmark.
  3. [Section 5.1 and Section 5.2] The decision to assign compositionality score 0 to all named entities by default, combined with the exclusion of definite-article constructions, can systematically affect the reported compositionality and domain-specificity distributions. For example, the higher ratio of non-compositional NCCs in Dúchas (0.36 vs 0.12) could partly reflect a different prevalence of named entities, which are scored 0 by construction rather than by semantic judgment. Please quantify how many of the reported non-compositional NCCs are named entities and provide an analysis with and without them.
  4. [Section 5.2] The statistics for the 270 UD-IDT NCCs and the 105 pilot NCCs are presented without an item list or per-annotator breakdown, so none of the averages, counts, or ratios can be checked. Please include the full list of NCCs with each annotator's scores, or clearly indicate that the dataset is available in a repository, and specify the number of sentences and annotators contributing to each statistic.
minor comments (5)
  1. [Section 4.2] The reference to agreement results appears as 'Table ??' and must be replaced with the actual table number.
  2. [Section 4.3] The definition of a two-word NCC is not fully precise: 'mí na meala' is excluded because a determiner is interleaved, but the example contains three orthographic words; please clarify whether 'two-word' refers to the two content words or to contiguous tokens.
  3. [Section 4.3] In the domain-specificity paragraph, the phrase 'and, is so is scored 1' should read 'and so is scored 1'.
  4. [Section 5.1] The bibliographic entry 'Christian-Brothers, 1999' is unusual; the author name should be formatted consistently with the rest of the reference list.
  5. [Section 6] Please state the license under which the annotations and the noun-compound collection will be released, given that both source corpora have their own licenses.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is an annotation resource paper whose conclusions rest on the collected annotations and inter-annotator agreement, not on any fitted parameter or self-derived prediction.

full rationale

The paper reports an annotation study of Irish noun compounds: it defines an NCC as a two-word contiguous noun-headed construction, collects ratings for compositionality, domain specificity, familiarity, and confidence, and analyzes the resulting distributions. There is no derivation chain in which an output is equivalent to an input by construction. The compositionality definition and the six-point scale are annotation conventions, not quantities fitted to the data and then re-predicted. Inter-annotator agreement (Cohen's weighted kappa, Table 1) is computed directly from the three annotators' scores and serves as an independent check on the reliability of the guidelines; no statistic is normalized or transformed into the quantity it is claimed to predict. Self-citations to McGuinness et al. (2020) and Lynn and Foster (2016) provide background on Irish UD annotation and the treebank source, but the paper's central contribution is the newly created pilot annotations and guidelines, and those citations do not force the reported compositionality or domain-specificity findings. The paper explicitly discloses its scope limitations (two-word NCs, exclusion of definite-article constructions and named entities), and while the absence of a release link is a reproducibility concern, it is not a circularity concern. No load-bearing step reduces to the paper's own inputs.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper introduces no fitted parameters or new entities. It relies on an operational definition of noun compounds and a standard agreement statistic.

assumptions (2)
  • domain assumption A noun compound is operationally defined as a contiguous two-word construction whose head is a noun and whose distribution is that of a noun, excluding constructions containing a definite article.
    Introduced in Section 4.3; drives what is and is not annotated, and the paper acknowledges it narrows the scope of the resource.
  • standard math Cohen's weighted kappa is an appropriate measure of agreement for the ordinal compositionality and domain-specificity ratings.
    Used in Section 4.2 (Table 1) without discussion; a common choice in annotation studies, but its interpretive thresholds are not stated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Annotating Compositionality Scores for Irish Noun Compounds is Hard Work." pith.science (2026). https://pith.science/paper/BKQMPEWL

@misc{pith2026250210061,
  author       = {Pith},
  title        = {Pith review of: Annotating Compositionality Scores for Irish Noun Compounds is Hard Work},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BKQMPEWL}},
  note         = {Machine review of arXiv:2502.10061}
}
read the original abstract

Noun compounds constitute a challenging construction for NLP applications, given their variability in idiomaticity and interpretation. In this paper, we present an analysis of compound nouns identified in Irish text of varied domains by expert annotators, focusing on compositionality as a key feature, but also domain specificity, as well as familiarity and confidence of the annotator giving the ratings. Our findings and the discussion that ensued contributes towards a greater understanding of how these constructions appear in Irish language, and how they might be treated separately from English noun compounds.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

36 extracted references · 19 canonical work pages

  1. [1]

    Timothy Baldwin and Su Nam Kim. 2010. https://api.semanticscholar.org/CorpusID:29511937 Multiword expressions . In Handbook of Natural Language Processing

  2. [4]

    Fillmore, Ralph Grishman, Nancy Ide, Alessandro Lenci, Catherine MacLeod, and Antonio Zampolli

    Nicoletta Calzolari, Charles J. Fillmore, Ralph Grishman, Nancy Ide, Alessandro Lenci, Catherine MacLeod, and Antonio Zampolli. 2002. http://www.lrec-conf.org/proceedings/lrec2002/pdf/259.pdf Towards best practice for multiword expressions in computational lexicons . In Proceedings of the Third International Conference on Language Resources and Evaluation...

  3. [5]

    Jiaao Chen, Xiaoman Pan, Dian Yu, Kaiqiang Song, Xiaoyang Wang, Dong Yu, and Jianshu Chen. 2023. http://arxiv.org/abs/2308.00304 Skills-in-context prompting: Unlocking compositionality in large language models

  4. [6]

    Christian-Brothers. 1999. Graim \'e ar Gaeilge na mBr \'a ithre Cr \' osta \' . An G \'u m, Baile \' A tha Cliath

  5. [8]

    Silvio Cordeiro, Aline Villavicencio, Marco Idiart, and Carlos Ramisch. 2019. https://doi.org/10.1162/coli_a_00341 Unsupervised compositionality prediction of nominal compounds . Computational Linguistics, 45(1):1--57

  6. [9]

    Meghdad Farahmand, Aaron Smith, and Joakim Nivre. 2015. https://doi.org/10.3115/v1/W15-0904 A multiword expression data set: Annotating non-compositionality and conventionalization for E nglish noun compounds . In Proceedings of the 11th Workshop on Multiword Expressions, pages 29--33, Denver, Colorado. Association for Computational Linguistics

  7. [10]

    Marcos Garcia, Tiago Kramer Vieira, Carolina Scarton, Marco Idiart, and Aline Villavicencio. 2021. https://doi.org/10.18653/v1/2021.acl-long.212 Assessing the representations of idiomaticity in vector models with a noun compound dataset labeled at type and token levels . In Proceedings of the 59th Annual Meeting of the Association for Computational Lingui...

  8. [11]

    Roxana Girju, Dan Moldovan, Marta Tatu, and Daniel Antohe. 2005. https://doi.org/10.1016/j.csl.2005.02.006 On the semantics of noun compounds . Comput. Speech Lang., 19(4):479–496

Show all 36 references
  1. [12]

    Jan-Christoph Klie, Michael Bugert, Beto Boullosa, Richard Eckart de Castilho, and Iryna Gurevych. 2018. http://tubiblio.ulb.tu-darmstadt.de/106270/ The inception platform: Machine-assisted and knowledge-oriented interactive annotation . In Proceedings of the 27th Internationa...

  2. [13]

    Teresa Lynn. 2022. Report on the I rish language. https://european-language-equality.eu/deliverables/. Technical Report D1.20, European Language Equality Project

  3. [14]

    Teresa Lynn and Jennifer Foster. 2016. Universal dependencies for irish. In Proceedings of the 2nd Celtic Language Technology Workshop, Paris, France

  4. [15]

    Sarah McGuinness, Jason Phelan, Abigail Walsh, and Teresa Lynn. 2020. https://aclanthology.org/2020.udw-1.15 Annotating MWE s in the I rish UD treebank . In Proceedings of the Fourth Workshop on Universal Dependencies (UDW 2020), pages 126--139, Barcelona, Spain (Online). Asso...

  5. [17]

    Nouvel, M

    D. Nouvel, M. Ehrmann, and S. Rosset. 2016. https://books.google.ie/books?id=2vpRCgAAQBAJ Named Entities for Computational Linguistics . Cognitive science series. Wiley

  6. [18]

    Digital Plan for the Irish Language Speech and Language Technologies 2023-2027

    Ailbhe Ní Chasaide, Neasa Ní Chiarán, Elaine Uí Dhonnchadha, Teresa Lynn, and John Judge. Digital Plan for the Irish Language Speech and Language Technologies 2023-2027 . Available at https://assets.gov.ie/241755/e82c256a-6f47-4ddb-8ce6-ff81df208bb1.pdf

  7. [20]

    Carlos Ramisch, Silvio Ricardo Cordeiro, Agata Savary, Veronika Vincze, Verginica Barbu Mititelu, Archna Bhatia, Maja Buljan, Marie Candito, Polona Gantar, Voula Giouli, Tunga G \"u ng \"o r, Abdelati Hawwari, Uxoa I \ n urrieta, Jolanta Kovalevskait \.e , Simon Krek, Timm Lic...

  8. [21]

    Carlos Ramisch, Agata Savary, Bruno Guillaume, Jakub Waszczuk, Marie Candito, Ashwini Vaidya, Verginica Barbu Mititelu, Archna Bhatia, Uxoa I \ n urrieta, Voula Giouli, Tunga G \"u ng \"o r, Menghan Jiang, Timm Lichte, Chaya Liebeskind, Johanna Monti, Renata Ramisch, Sara Stym...

  9. [22]

    Siva Reddy, Diana McCarthy, and Suresh Manandhar. 2011. https://aclanthology.org/I11-1024 An empirical study on compositionality in compound nouns . In Proceedings of 5th International Joint Conference on Natural Language Processing, pages 210--218, Chiang Mai, Thailand. Asian...

  10. [23]

    Ivan Sag, Timothy Baldwin, Francis Bond, Ann Copestake, and Dan Flickinger. 2002. https://doi.org/10.1007/3-540-45715-1_1 Multiword expressions: A pain in the neck for nlp . pages 1--15

  11. [24]

    Agata Savary, Carlos Ramisch, Silvio Cordeiro, Federico Sangati, Veronika Vincze, Behrang QasemiZadeh, Marie Candito, Fabienne Cap, Voula Giouli, Ivelina Stoyanova, and Antoine Doucet. 2017. https://doi.org/10.18653/v1/W17-1704 The PARSEME shared task on automatic identificati...

  12. [26]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked islrn pid label extra.label sort.label short.list INTEGERS output.st...

  13. [27]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

  14. [28]

    Laurie Bauer. 2019. https://doi.org/10.1515/9783110632446-002 Compounds and multi-word expressions in English . Complex Lexical Units: Compounds and Multi-Word Expressions, pages 45--68

  15. [29]

    Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2019. https://arxiv.org/abs/1911.11641 PIQA: Reasoning about Physical Commonsense in Natural Language . Preprint, arXiv:1911.11641

  16. [30]

    Collins, Catherine Wong, Jiahai Feng, Megan Wei, and Joshua B

    Katherine M. Collins, Catherine Wong, Jiahai Feng, Megan Wei, and Joshua B. Tenenbaum. 2022. https://arxiv.org/abs/2205.05718 Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning ta...

  17. [31]

    Teresa Lynn and Jennifer Foster. 2016. Universal Dependencies for Irish . In Proceedings of the 2nd Celtic Language Technology Workshop, Paris, France

  18. [32]

    Teresa Lynn, Jennifer Foster, Sarah McGuinness, Abigail Walsh, Jason Phelan, and Kevin Scannell. 2023. Universal Dependencies Irish Dependency Treebank (v2.12)

  19. [33]

    Kanishka Misra, Julia Taylor Rayz, and Allyson Ettinger. 2023. https://arxiv.org/abs/2210.01963 Comps: Conceptual minimal pair sentences for testing robust property knowledge and its inheritance in pre-trained language models . Preprint, arXiv:2210.01963

  20. [34]

    Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer

    Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. https://arxiv.org/abs/1802.05365 Deep contextualized word representations

  21. [35]

    The Dúchas Project . 2016. National folklore collection: The schools’ collection (digitized)

  22. [36]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. https://arxiv.org/abs/1706.03762 Attention is all you need . Preprint, arXiv:1706.03762

  23. [37]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  24. [38]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

  25. [39]

    Lynn, Teresa and Foster, Jennifer and McGuinness, Sarah and Walsh, Abigail and Phelan, Jason and Scannell, Kevin. 2023. Universal Dependencies Irish Dependency Treebank (v2.12). Universal Dependencies. LINDAT/CLARIN

  26. [40]

    The Dúchas Project. 2016. National Folklore Collection: The Schools’ Collection (digitized). The Dúchas Project. Hosted at https://www.duchas.ie/en

  27. [41]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked islrn label extra.label sort.label short.list INTEGERS output.state ...

  28. [42]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.