REVIEW 4 major objections 5 minor 36 references
Annotating Compositionality Scores for Irish Noun Compounds is Hard Work
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper creates the first collection of Irish noun compounds with compositionality ratings and argues it is a reliable basis for Irish-specific noun-compound resources and for evaluating language models in Irish.
desk verdict First Irish noun-compound compositionality annotation effort, honestly reported, but the claimed dataset is not actually accessible from the paper. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The operative object is the noun compound candidate (NCC), defined in the guidelines as a contiguous two-word construction in which one component is a noun or an adjective dependent on a noun head, and the whole construction has nominal distribution. Interleaved determiners disqualify a candidate; named entities can be annotated but receive compositionality 0 by default. The measuring instruments are the six-point compositionality scale (0 opaque to 5 transparent) and the three-point domain-specificity, familiarity, and confidence scales. The pilot-task refinement loop—three rounds with pair-wise weighted kappa values mostly between 0.3 and 0.64—is the procedural machinery that turned initial disagreement into a usable guideline document.
What would settle it
The reported folklore-versus-treebank contrast would be settled by re-annotating the same sentences under a wider definition that includes interleaved-determiner constructions such as 'mí na meala' (honeymoon); if the opaque-compound ratio moves substantially toward the folklore figure, the operational definition is producing the observed difference.
Extended reading notes
Core claim
The paper's central claim is that noun-compound annotation for Irish is tractable and informative once the object is defined narrowly enough and the guidelines are shaped by pilot disagreements. Working from a definition of a noun compound as a contiguous two-word phrase whose head is a noun and whose distribution is nominal, the authors show that expert annotators can assign six-point compositionality scores together with domain-specificity, familiarity, and confidence ratings, and that the resulting judgements reveal a systematic contrast between text types: the dialectal folklore corpus contains a higher share of opaque compounds (average ratio 0.36 versus 0.12 in the modern treebank) and receives lower domain-specificity and confidence scores. They also report that translating English noun compounds into Irish rarely produces a two-word Irish compound, since many English compounds map to single Irish words or to possessive constructions with interleaved articles. The released pilot annotations and the associated guidelines are offered as the first collection of Irish noun compounds with compositionality ratings.
Load-bearing premise
The load-bearing premise is that defining an Irish noun compound as a contiguous two-word phrase headed by a noun, with interleaved determiners and named entities excluded or scored as opaque, does not bias the measured compositionality away from the constructions Irish speakers actually use.
Editorial extensions
If this is right
- If the pilot corpus and guidelines are reliable, they provide the first Irish-language benchmark for testing whether language models' handling of idiomatic compounds transfers beyond English.
- The observed contrast between folklore and modern treebank data implies that any Irish noun-compound resource must be domain-balanced, because folklore-style text over-represents opaque compounds while modern text over-represents compositional ones.
- The translation finding—only 30 of 280 English noun compounds map to two-word Irish compounds—implies that English compound lexicons cannot simply be projected onto Irish; Irish resources have to be built from Irish text.
- The average inter-annotator agreement around 0.5 weighted kappa implies that compositionality scores should be accompanied by confidence and familiarity metadata, and that downstream uses should treat the scores as graded rather than categorical.
- The deliberate exclusion of interleaved-determiner constructions narrows the current resource but also marks the boundary for a natural future extension.
Reading between the lines
- Editorial inference: the operational definition (contiguous two words, noun head, no interleaved determiner) likely under-counts possessive and article-bearing Irish constructions that carry non-compositional meaning, so the released statistics may understate the proportion of opaque expressions in ordinary Irish text.
- Editorial inference: a direct testable next step is to feed the released noun compound candidates to multilingual language models and compare their compositionality judgements with the human scores, which would reveal whether Irish opacity is harder for models than the English and Portuguese cases studied elsewhere.
- Editorial inference: the translation result points to a linguistic asymmetry—English tends to pack non-compositional meanings into noun-plus-noun compounds, while Irish tends to realize them as single words or possessive phrases—which would change how cross-lingual multiword-expression identification should be designed.
- Editorial inference: the kappa pattern across pilots suggests that annotator training on hard cases matters more than the choice of rating scale, so a small follow-up study varying annotator background and guideline examples could test whether the guidelines transfer beyond the original annotator pool.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports ongoing annotation work for Irish noun compounds, presenting guidelines for identifying noun compound candidates (NCCs) and for scoring compositionality, domain specificity, annotator familiarity, and confidence. The authors describe three pilot tasks with expert annotators, report Cohen's weighted kappa agreement scores, discuss difficult cases such as definite-article constructions and named entities, and give preliminary statistics contrasting the Dúchas folklore corpus with the UD-IDT treebank. The paper claims to provide the first collection of Irish noun compounds with compositionality ratings and to make pilot annotations publicly available.
Significance. If the dataset and guidelines were released, the work would be a useful first resource for Irish multiword-expression research and for evaluating language models on Irish noun compounds. The use of multiple expert annotators with different dialect backgrounds and the explicit discussion of annotation difficulties are valuable. The paper is honest about its scope limitations, such as restricting NCCs to two-word contiguous constructions and excluding definite-article constructions. However, the central contribution is a resource that is not currently available, and the inter-annotator agreement is moderate, so the empirical claims cannot yet be fully verified.
major comments (4)
- [Section 1 and Section 6] The paper states in Section 1 that 'the pilot task annotations are made available for public use', but no URL, DOI, repository, or appendix is provided; Section 6 instead says the collection 'will be released alongside the annotated corpus'. Since the central claim is the creation of a first-of-its-kind Irish noun-compound resource with compositionality ratings, the absence of any accessible artifact makes the core empirical claims unverifiable. The final version must include a working link or DOI and ideally the full item-level annotations.
- [Table 1 and Section 4.2] The weighted kappa values in Table 1 range from 0.30 to 0.64, with many pairs below 0.55, which is generally considered only moderate or weak agreement. The text in Section 4.2 says the agreement results 'provide a glimpse into the level of consensus', but the later framing that the guidelines support 'reliable annotation' is not supported by these numbers. Please report the weighting scheme used, confidence intervals, and the number of items per pilot, and discuss whether the pilot tasks were used iteratively to revise the guidelines or as a final reliability benchmark.
- [Section 5.1 and Section 5.2] The decision to assign compositionality score 0 to all named entities by default, combined with the exclusion of definite-article constructions, can systematically affect the reported compositionality and domain-specificity distributions. For example, the higher ratio of non-compositional NCCs in Dúchas (0.36 vs 0.12) could partly reflect a different prevalence of named entities, which are scored 0 by construction rather than by semantic judgment. Please quantify how many of the reported non-compositional NCCs are named entities and provide an analysis with and without them.
- [Section 5.2] The statistics for the 270 UD-IDT NCCs and the 105 pilot NCCs are presented without an item list or per-annotator breakdown, so none of the averages, counts, or ratios can be checked. Please include the full list of NCCs with each annotator's scores, or clearly indicate that the dataset is available in a repository, and specify the number of sentences and annotators contributing to each statistic.
minor comments (5)
- [Section 4.2] The reference to agreement results appears as 'Table ??' and must be replaced with the actual table number.
- [Section 4.3] The definition of a two-word NCC is not fully precise: 'mí na meala' is excluded because a determiner is interleaved, but the example contains three orthographic words; please clarify whether 'two-word' refers to the two content words or to contiguous tokens.
- [Section 4.3] In the domain-specificity paragraph, the phrase 'and, is so is scored 1' should read 'and so is scored 1'.
- [Section 5.1] The bibliographic entry 'Christian-Brothers, 1999' is unusual; the author name should be formatted consistently with the rest of the reference list.
- [Section 6] Please state the license under which the annotations and the noun-compound collection will be released, given that both source corpora have their own licenses.
Circularity Check
No circularity: the paper is an annotation resource paper whose conclusions rest on the collected annotations and inter-annotator agreement, not on any fitted parameter or self-derived prediction.
full rationale
The paper reports an annotation study of Irish noun compounds: it defines an NCC as a two-word contiguous noun-headed construction, collects ratings for compositionality, domain specificity, familiarity, and confidence, and analyzes the resulting distributions. There is no derivation chain in which an output is equivalent to an input by construction. The compositionality definition and the six-point scale are annotation conventions, not quantities fitted to the data and then re-predicted. Inter-annotator agreement (Cohen's weighted kappa, Table 1) is computed directly from the three annotators' scores and serves as an independent check on the reliability of the guidelines; no statistic is normalized or transformed into the quantity it is claimed to predict. Self-citations to McGuinness et al. (2020) and Lynn and Foster (2016) provide background on Irish UD annotation and the treebank source, but the paper's central contribution is the newly created pilot annotations and guidelines, and those citations do not force the reported compositionality or domain-specificity findings. The paper explicitly discloses its scope limitations (two-word NCs, exclusion of definite-article constructions and named entities), and while the absence of a release link is a reproducibility concern, it is not a circularity concern. No load-bearing step reduces to the paper's own inputs.
Assumptions & free parameters
assumptions (2)
- domain assumption A noun compound is operationally defined as a contiguous two-word construction whose head is a noun and whose distribution is that of a noun, excluding constructions containing a definite article.
- standard math Cohen's weighted kappa is an appropriate measure of agreement for the ordinal compositionality and domain-specificity ratings.
Cite this review
Pith. "Pith review of Annotating Compositionality Scores for Irish Noun Compounds is Hard Work." pith.science (2026). https://pith.science/paper/BKQMPEWL
@misc{pith2026250210061,
author = {Pith},
title = {Pith review of: Annotating Compositionality Scores for Irish Noun Compounds is Hard Work},
year = {2026},
howpublished = {\url{https://pith.science/paper/BKQMPEWL}},
note = {Machine review of arXiv:2502.10061}
}
read the original abstract
Noun compounds constitute a challenging construction for NLP applications, given their variability in idiomaticity and interpretation. In this paper, we present an analysis of compound nouns identified in Irish text of varied domains by expert annotators, focusing on compositionality as a key feature, but also domain specificity, as well as familiarity and confidence of the annotator giving the ratings. Our findings and the discussion that ensued contributes towards a greater understanding of how these constructions appear in Irish language, and how they might be treated separately from English noun compounds.
Reference graph
Works this paper leans on
-
[1]
Timothy Baldwin and Su Nam Kim. 2010. https://api.semanticscholar.org/CorpusID:29511937 Multiword expressions . In Handbook of Natural Language Processing
work page 2010
-
[4]
Fillmore, Ralph Grishman, Nancy Ide, Alessandro Lenci, Catherine MacLeod, and Antonio Zampolli
Nicoletta Calzolari, Charles J. Fillmore, Ralph Grishman, Nancy Ide, Alessandro Lenci, Catherine MacLeod, and Antonio Zampolli. 2002. http://www.lrec-conf.org/proceedings/lrec2002/pdf/259.pdf Towards best practice for multiword expressions in computational lexicons . In Proceedings of the Third International Conference on Language Resources and Evaluation...
work page 2002
-
[5]
Jiaao Chen, Xiaoman Pan, Dian Yu, Kaiqiang Song, Xiaoyang Wang, Dong Yu, and Jianshu Chen. 2023. http://arxiv.org/abs/2308.00304 Skills-in-context prompting: Unlocking compositionality in large language models
arXiv 2023
-
[6]
Christian-Brothers. 1999. Graim \'e ar Gaeilge na mBr \'a ithre Cr \' osta \' . An G \'u m, Baile \' A tha Cliath
work page 1999
-
[8]
Silvio Cordeiro, Aline Villavicencio, Marco Idiart, and Carlos Ramisch. 2019. https://doi.org/10.1162/coli_a_00341 Unsupervised compositionality prediction of nominal compounds . Computational Linguistics, 45(1):1--57
-
[9]
Meghdad Farahmand, Aaron Smith, and Joakim Nivre. 2015. https://doi.org/10.3115/v1/W15-0904 A multiword expression data set: Annotating non-compositionality and conventionalization for E nglish noun compounds . In Proceedings of the 11th Workshop on Multiword Expressions, pages 29--33, Denver, Colorado. Association for Computational Linguistics
-
[10]
Marcos Garcia, Tiago Kramer Vieira, Carolina Scarton, Marco Idiart, and Aline Villavicencio. 2021. https://doi.org/10.18653/v1/2021.acl-long.212 Assessing the representations of idiomaticity in vector models with a noun compound dataset labeled at type and token levels . In Proceedings of the 59th Annual Meeting of the Association for Computational Lingui...
-
[11]
Roxana Girju, Dan Moldovan, Marta Tatu, and Daniel Antohe. 2005. https://doi.org/10.1016/j.csl.2005.02.006 On the semantics of noun compounds . Comput. Speech Lang., 19(4):479–496
Show all 36 references
-
[12]
Jan-Christoph Klie, Michael Bugert, Beto Boullosa, Richard Eckart de Castilho, and Iryna Gurevych. 2018. http://tubiblio.ulb.tu-darmstadt.de/106270/ The inception platform: Machine-assisted and knowledge-oriented interactive annotation . In Proceedings of the 27th Internationa...
2018
-
[13]
Teresa Lynn. 2022. Report on the I rish language. https://european-language-equality.eu/deliverables/. Technical Report D1.20, European Language Equality Project
2022
-
[14]
Teresa Lynn and Jennifer Foster. 2016. Universal dependencies for irish. In Proceedings of the 2nd Celtic Language Technology Workshop, Paris, France
2016
-
[15]
Sarah McGuinness, Jason Phelan, Abigail Walsh, and Teresa Lynn. 2020. https://aclanthology.org/2020.udw-1.15 Annotating MWE s in the I rish UD treebank . In Proceedings of the Fourth Workshop on Universal Dependencies (UDW 2020), pages 126--139, Barcelona, Spain (Online). Asso...
2020
-
[17]
Nouvel, M
D. Nouvel, M. Ehrmann, and S. Rosset. 2016. https://books.google.ie/books?id=2vpRCgAAQBAJ Named Entities for Computational Linguistics . Cognitive science series. Wiley
2016
-
[18]
Digital Plan for the Irish Language Speech and Language Technologies 2023-2027
Ailbhe Ní Chasaide, Neasa Ní Chiarán, Elaine Uí Dhonnchadha, Teresa Lynn, and John Judge. Digital Plan for the Irish Language Speech and Language Technologies 2023-2027 . Available at https://assets.gov.ie/241755/e82c256a-6f47-4ddb-8ce6-ff81df208bb1.pdf
2023
-
[20]
Carlos Ramisch, Silvio Ricardo Cordeiro, Agata Savary, Veronika Vincze, Verginica Barbu Mititelu, Archna Bhatia, Maja Buljan, Marie Candito, Polona Gantar, Voula Giouli, Tunga G \"u ng \"o r, Abdelati Hawwari, Uxoa I \ n urrieta, Jolanta Kovalevskait \.e , Simon Krek, Timm Lic...
2018
-
[21]
Carlos Ramisch, Agata Savary, Bruno Guillaume, Jakub Waszczuk, Marie Candito, Ashwini Vaidya, Verginica Barbu Mititelu, Archna Bhatia, Uxoa I \ n urrieta, Voula Giouli, Tunga G \"u ng \"o r, Menghan Jiang, Timm Lichte, Chaya Liebeskind, Johanna Monti, Renata Ramisch, Sara Stym...
2020
-
[22]
Siva Reddy, Diana McCarthy, and Suresh Manandhar. 2011. https://aclanthology.org/I11-1024 An empirical study on compositionality in compound nouns . In Proceedings of 5th International Joint Conference on Natural Language Processing, pages 210--218, Chiang Mai, Thailand. Asian...
2011
-
[23]
Ivan Sag, Timothy Baldwin, Francis Bond, Ann Copestake, and Dan Flickinger. 2002. https://doi.org/10.1007/3-540-45715-1_1 Multiword expressions: A pain in the neck for nlp . pages 1--15
2002 doi
-
[24]
Agata Savary, Carlos Ramisch, Silvio Cordeiro, Federico Sangati, Veronika Vincze, Behrang QasemiZadeh, Marie Candito, Fabienne Cap, Voula Giouli, Ivelina Stoyanova, and Antoine Doucet. 2017. https://doi.org/10.18653/v1/W17-1704 The PARSEME shared task on automatic identificati...
2017 doi
-
[26]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked islrn pid label extra.label sort.label short.list INTEGERS output.st...
-
[27]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
-
[28]
Laurie Bauer. 2019. https://doi.org/10.1515/9783110632446-002 Compounds and multi-word expressions in English . Complex Lexical Units: Compounds and Multi-Word Expressions, pages 45--68
2019 doi
-
[29]
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2019. https://arxiv.org/abs/1911.11641 PIQA: Reasoning about Physical Commonsense in Natural Language . Preprint, arXiv:1911.11641
2019 arXiv
-
[30]
Collins, Catherine Wong, Jiahai Feng, Megan Wei, and Joshua B
Katherine M. Collins, Catherine Wong, Jiahai Feng, Megan Wei, and Joshua B. Tenenbaum. 2022. https://arxiv.org/abs/2205.05718 Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning ta...
2022 arXiv
-
[31]
Teresa Lynn and Jennifer Foster. 2016. Universal Dependencies for Irish . In Proceedings of the 2nd Celtic Language Technology Workshop, Paris, France
2016
-
[32]
Teresa Lynn, Jennifer Foster, Sarah McGuinness, Abigail Walsh, Jason Phelan, and Kevin Scannell. 2023. Universal Dependencies Irish Dependency Treebank (v2.12)
2023
-
[33]
Kanishka Misra, Julia Taylor Rayz, and Allyson Ettinger. 2023. https://arxiv.org/abs/2210.01963 Comps: Conceptual minimal pair sentences for testing robust property knowledge and its inheritance in pre-trained language models . Preprint, arXiv:2210.01963
2023 arXiv
-
[34]
Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. https://arxiv.org/abs/1802.05365 Deep contextualized word representations
2018 arXiv
-
[35]
The Dúchas Project . 2016. National folklore collection: The schools’ collection (digitized)
2016
-
[36]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. https://arxiv.org/abs/1706.03762 Attention is all you need . Preprint, arXiv:1706.03762
2017 arXiv
-
[37]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[38]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
-
[39]
Lynn, Teresa and Foster, Jennifer and McGuinness, Sarah and Walsh, Abigail and Phelan, Jason and Scannell, Kevin. 2023. Universal Dependencies Irish Dependency Treebank (v2.12). Universal Dependencies. LINDAT/CLARIN
2023
-
[40]
The Dúchas Project. 2016. National Folklore Collection: The Schools’ Collection (digitized). The Dúchas Project. Hosted at https://www.duchas.ie/en
2016
-
[41]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked islrn label extra.label sort.label short.list INTEGERS output.state ...
-
[42]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.