{"id":"51cc9b4e-be13-46fd-a3c2-3eb8da96a809","arxiv_id":"2508.11828","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of 53 idiom datasets finds that psycholinguistic norming resources and computational idiom corpora have essentially no overlap or cross-usage.","lead":"This paper reviews 53 public and private datasets used to study idioms in psychology and in computer language processing, comparing what each field measures. It finds that the two research communities barely use each other's data, and argues for shared metadata and dataset formats.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central 'no relation' claim is asserted without measuring relation and is contradicted by the survey's own entries (e.g., Peng et al. 2014; Senaldi 2019).","rationale":"The reader's weakest assumption focuses on corpus completeness. While that is a legitimate concern for any negative claim, the more immediate problem is that the paper's own data contain apparent counterexamples to 'no relation.' The survey does not operationalize what 'relation' means, so the claim is unfalsifiable as stated. The internal contradictions (Peng et al. 2014 using valence ratings; Senaldi 2019 bridging both fields; several computational datasets using compositionality ratings) mean the central claim is unsupported regardless of whether the 53-dataset sample is exhaustive. The paper is still valuable as a descriptive catalog, but the synthesizing conclusion must be either backed by a quantitative relational analysis or heavily qualified. Since the reader's verdict CONDITIONAL already permits revisions, I keep that verdict; my concern is different but leads to the same required revision.","tokens_in":12330,"tokens_out":5987,"duration_ms":66320,"concrete_test":"Build a citation-reuse graph from the paper's 53 listed datasets: for each dataset, identify whether it is cited by or used as a gold standard in any paper from the other community (e.g., search Google Scholar for 'PANIG' within computational idiom papers, and 'MAGPIE' within psycholinguistic norming papers). If at least one cross-citation or dataset reuse exists, the 'no relation' claim is falsified as stated. A minimal check: determine whether Peng et al. (2014) is cited in Section 2's discussion and whether Senaldi (2019) appears in both fields' references; if so, the survey already documents a relation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central claim—'there seems to be no relation yet between psycholinguistic and computational research on idioms' (Section 3.5)—is a negative existential that the paper does not substantiate. Nowhere is 'relation' operationally defined; no citation analysis, author-overlap check, dataset-reuse analysis, or shared-task analysis is provided. More importantly, the paper's own sections contradict the claim. Section 2 says Peng et al. (2014) 'used the rated valence of idiom component words for automated detection of idioms in texts'—a computational method built on a psycholinguistic rating dimension. Table 2 includes Senaldi (2019), a thesis explicitly titled 'Working both sides of the street: computational and psycholinguistic investigations on idiomatic variability.' Several computational datasets (Reddy et al. 2011; Cordeiro et al. 2019; NCS; Swedish MWEs) use human compositionality ratings, a psycholinguistic construct. Thus, either 'relation' is defined so narrowly that the claim is trivial, or the claim is false. The type/token distinction does not rescue it: many computational datasets (MAGPIE, SLIDE, Idiom Paraphrases) contain hundreds/thousands of idiom types, while psycholinguistic norms could be applied to tokens. The central conclusion therefore needs either a quantitative relational analysis or a substantial hedge.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper surveys 53 datasets for idiom research, split between psycholinguistic norming resources (Table 1) and computational-linguistics resources (Table 2), and summarizes their language coverage, annotation dimensions, task framings, sizes, and availability. It identifies heterogeneity in annotation schemes and argues that the two research traditions have remained largely unconnected, attributing this to a type/token divide: psycholinguistic datasets treat idioms as types, while computational datasets work with tokens in context. The paper also proposes possible future bridges, such as sentiment/valence norms and computational modeling of human ratings.","tokens_in":12610,"tokens_out":3395,"duration_ms":40080,"significance":"If its descriptive content is accurate, this survey is a useful reference for researchers seeking idiom datasets: the appendix tables consolidate a scattered literature, the GitHub resource list is a practical asset, and spot-checks of the reported statistics (e.g., MAGPIE, VNC-Tokens, RU Idioms, LIdioms, PANIG, Gavilán et al.) match the cited publications. The organizing distinction between psycholinguistic type-level norms and computational token-level resources is a suggestive framing. However, the paper's central synthesizing claim—that there is 'no relation yet' between the two fields—is asserted without operationalization or evidence and is in tension with several entries in the paper's own tables. This weakens the main conclusion, although the descriptive survey remains valuable as a reference work.","major_comments":[{"comment":"The central claim that 'there seems to be no relation yet between psycholinguistic and computational research on idioms' is a negative existential that is never operationalized and no evidence is provided for it: there is no citation-overlap analysis, dataset-reuse analysis, author-overlap check, or shared-task analysis. Moreover, the paper's own contents contradict the claim as stated. Section 2 explicitly notes that Peng et al. (2014) used psycholinguistic valence ratings of idiom component words for automated idiom detection. Table 2 includes Senaldi (2019), a thesis explicitly titled 'Working both sides of the street: computational and psycholinguistic investigations on idiomatic variability.' Several computational datasets in Table 2 (Reddy et al. 2011; Cordeiro et al. 2019; NCS; Swedish MWEs) rely on human compositionality ratings—a psycholinguistic construct. The claim needs eithe","section":"§3.5"},{"comment":"The type/token distinction offered as the explanation for the alleged lack of relation is not supported by Tables 1 and 2. Many computational datasets are explicitly type-level resources: IDIOMENT (580 idiom types), SLIDE (5,000 idiom types), LIdioms (815 types), CCT (7,395 types), CIKB (38K types), ChID (3,848 types), and IdiomKB. Conversely, psycholinguistic norming studies could in principle be applied to tokens, and some psycholinguistic studies use multiple instances or contexts. Thus the type/token dichotomy is not a clean structural barrier, and it cannot bear the weight of the no-relation conclusion.","section":"§3.5"},{"comment":"The survey gives no search protocol, inclusion/exclusion criteria, date range, or database/source list. Section 1 only says the search was 'extensive,' and the Limitations section concedes 'some resources might have been missed.' For a negative existential claim about the field, the representativeness of the 53 selected datasets is load-bearing: a single missed bridging dataset would weaken the conclusion. The authors should either describe a systematic, reproducible selection methodology or weaken the conclusion to a more observational statement about the datasets they actually surveyed.","section":"§1, Limitations, Appendix"},{"comment":"The survey includes nominal-compound compositionality resources (Reddy et al. 2011; Cordeiro et al. 2019; NCS; Swedish MWEs) as 'idiom datasets,' but this classification is not justified. Nominal compounds are multword expressions, not necessarily idioms; including them broadens the scope beyond the abstract's definition of idiomatic expressions whose meanings cannot be inferred from their parts. The paper should either defend this inclusion or exclude such resources, since the dataset count and the trends derived from it depend on this scope decision.","section":"Tables 1–2, scope"}],"minor_comments":[{"comment":"The row for Senaldi (2019) lists only '90 verb-noun and 24 adjective-noun expressions (types)' without indicating the task, rating dimensions, or the thesis's explicit computational+psycholinguistic design. Adding one or two descriptors would make the table more informative.","section":"Table 2, Senaldi 2019 row"},{"comment":"The entry says '5.4K instances of 100 idiomatic expressions (3K literal, 2.4K idiomatic)'—should specify '100 idiom types' rather than 'expressions' for consistency with the rest of the table.","section":"Table 2, RU Idioms row"},{"comment":"Missing comma/punctuation: '10K idioms (types) 262781 sentences' should read '10K idioms (types), 262,781 sentences'.","section":"Table 2, ID10M row"},{"comment":"Typo: 'probe LLMs inference' should be 'probe LLMs' inference'.","section":"§3.4"},{"comment":"The Beck (2020) row has no count; if the thesis includes norming data for a specific number of English idioms, that number should be reported, or the row should state that no norming count is specified.","section":"Table 1, Beck 2020 row"},{"comment":"The availability column for Morid and Sabourin (2024) is marked '-' in Table 1; the text says 'most of those datasets are publicly available.' Clarify whether this resource is unavailable or whether the dash denotes something else.","section":"§2, Morid and Sabourin"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a survey paper, not a methods paper, and judged as a survey it's useful. What's new: it pulls together 53 idiom datasets from psycholinguistics and computational linguistics, organizes them by discipline and task, and offers a type/token framing for why the two communities don't share data. The dataset tables are a genuinely handy reference; I spot-checked MAGPIE, VNC-Tokens, RU Idioms, LIdioms, PANIG, and Gavilán et al. against the originals and the numbers match.\n\nThe soft spot is the headline finding. Section 3.5 says 'there seems to be no relation yet between psycholinguistic and computational research on idioms,' but relation is never operationalized, and the paper's own entries undercut it. Peng et al. (2014) is a computational idiom-detection model built on psycholinguistic valence ratings; Senaldi (2019) is a dissertation explicitly straddling both sides; several computational datasets (Reddy et al., Cordeiro et al., NCS, Swedish MWEs) use human compositionality ratings, which is a psycholinguistic construct. So the claim either needs a narrow definition of 'relation' that would make it trivially true, or it needs a quantitative relational analysis — citation overlap, author overlap, dataset reuse — to hold up. The type/token distinction is suggestive but not a proof of disconnection.\n\nThe other issue is that the survey gives no search protocol, inclusion criteria, or date range, and its own Limitations section concedes resources may have been missed. For a descriptive survey that's a minor transparency problem, but it matters because the central negative claim depends on the 53 datasets being representative.\n\nThat said, the paper is honest: it flags the construct-comparability issue in Section 2, and the Limitations are stated rather than buried. The descriptive core is accurate and useful. If I were the editor, I'd send it to a referee — it deserves a careful reader who can push on the no-relation claim and ask for either a hedge or an actual relation analysis. With that revision, it becomes a solid reference piece for anyone working on idioms, whether they come from the norming side or the NLP side.","headline":"Useful reference map of 53 idiom datasets, but the central 'no relation' claim is asserted without evidence and undercut by the paper's own entries.","tokens_in":13124,"tokens_out":1734,"would_cite":true,"duration_ms":17865,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey of 53 idiom datasets finds no bridge between two research traditions.","keywords":["idiom datasets","psycholinguistic norms","computational linguistics","figurative language","multiword expressions","type-token distinction","survey"],"falsifier":"Search the reference lists of the surveyed computational idiom papers for citations to psycholinguistic norming studies (or vice versa). Finding even one dataset built on both traditions — for example, a computational idiomaticity corpus whose labels were derived from psycholinguistic norming ratings, or a norming study that used a computational idiom corpus to select items — would weaken the paper's claim that the two fields have no relation.","tokens_in":12150,"feed_emoji":"📚","tokens_out":6147,"duration_ms":62336,"temperature":0.7,"pith_summary":"The paper surveys 53 idiom datasets used in psycholinguistics and computational linguistics, cataloging their content, form, and intended use. Psycholinguistic resources norm idioms as types, providing aggregated human ratings on dimensions like familiarity, transparency, and compositionality; computational datasets work with idioms as tokens, labeling individual instances in context for tasks like idiomaticity detection, disambiguation, paraphrasing, and multilingual modeling. The survey's central claim is that these two traditions have not yet connected: there is no dataset, benchmark, or shared annotation scheme linking psycholinguistic norms to computational idiom resources. The authors argue this gap is structural rather than technical, rooted in the type/token divide, and suggest two possible convergence points: valence/sentiment of idioms, and using computational methods to model human ratings. If correct, the paper implies that progress in idiom-aware NLP will depend on deliberately building bridges between the two fields.","feed_headline":"Idiom research's two traditions never met, survey of 53 datasets shows","feed_subtitle":"Psycholinguists norm idioms as types; computationalists label them as tokens — the gap is structural, not technical.","key_machinery":"The central object is the survey's comparative framework: a two-table inventory of 53 idiom datasets organized by discipline (psycholinguistic norming studies vs computational corpora/benchmarks). The analytical work is done by grouping datasets along three axes — content (rating dimensions vs task labels), form (types vs tokens), and intended use (experimental control vs NLP evaluation). The type/token distinction is the load-bearing identity: it explains why annotation schemes do not transfer across the two traditions and why no unified benchmark exists.","core_discovery":"The central discovery is a gap, established by systematic comparison: the paper catalogs 53 idiom datasets and finds that psycholinguistic norming studies and computational idiom corpora operate in parallel, with no dataset, benchmark, or shared annotation scheme connecting them. Psycholinguistic resources aggregate Likert ratings of idiom properties for idioms treated as types; computational resources annotate instances of idioms in textual context, treating them as tokens. The authors identify the type/token distinction as the structural barrier: because the two traditions frame idioms at different granularities, their data cannot be directly compared or combined. They propose two possible","pith_inferences":["A direct test of the 'no relation' claim would be a citation analysis: count cross-citations between psycholinguistic idiom norming papers and computational idiom papers; if the paper is right, such cross-citations are rare.","The type/token gap may be a special case of a broader phenomenon in multiword-expression research, where lexicographic/psycholinguistic resources describe expressions while NLP systems need surface instances; the same split appears in compounds and other multiword units.","The survey's list suggests a concrete missing resource: no dataset yet provides both normed psycholinguistic ratings and per-instance contextual annotations for the same idiom set; building one would directly test the paper's central gap.","If the paper's diagnosis is correct, embedding norming dimensions as auxiliary labels in token-level datasets could improve idiom representation in language models, though this remains untested."],"forward_implications":["Future idiom benchmarks should be built with explicit metadata about idiom types vs tokens, and rating dimensions should be defined consistently across norming studies.","Psycholinguistic norms (familiarity, age of acquisition, valence) could be attached to computational instance-level datasets to test whether model errors track human-rated properties.","Computational methods could be used to predict or explain human ratings of idiom properties, one of the convergence paths the paper names.","Multilingual idiom datasets (LIdioms, IMIL, PETCI, SemEval-2022) are growing, but current resources lack semantically aligned idiom instances across languages, limiting cross-lingual bridge-building.","Shared metadata conventions and interoperable formats would help integration across the two traditions."],"supporting_citations":[{"why":"Provides the VNC-Tokens dataset, a canonical instance-level (token-based) computational resource for idiomaticity detection.","marker":"Fazly et al., 2009"},{"why":"Exemplifies type-based psycholinguistic norming, with ratings for 245 Italian idioms on familiarity, AoA, and compositionality.","marker":"Tabossi et al., 2011"},{"why":"Introduces MAGPIE, a large token-level corpus of 56,622 idiom instances in context, representing the computational token paradigm.","marker":"Haagsma et al., 2020"},{"why":"Provides PANIG, a German psycholinguistic norming study with affective ratings, illustrating the type-based approach and the valence dimension.","marker":"Citron et al., 2016"},{"why":"Supplies the theoretical distinctions (compositionality, transparency, analyzability) that many norming studies operationalize.","marker":"Nunberg et al., 1994"},{"why":"Describes SemEval-2013 Task 5b, a benchmark for idiom disambiguation with instance-level annotations, a key computational resource.","marker":"Korkontzelos et al., 2013"},{"why":"Presents the IDIX corpus, an instance-level idiomaticity corpus that anchors the token-based computational approach.","marker":"Sporleder et al., 2010"},{"why":"Documents SemEval-2022 Task 2, extending instance-level idiomaticity detection to three languages, showing the computational field's multilingual push.","marker":"Tayyar Madabushi et al., 2022"}],"fun_headline_variants":["53 idiom datasets, zero bridges: type/token gap divides fields","Idiom studies: two silos, 53 datasets, no connection","Survey finds psycholinguistic and computational idiom data don't mix","Idiom datasets: psycholinguists and computationalists speak different languages","Type vs token: why idiom research has two incompatible worlds"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The conclusion that psycholinguistic and computational idiom research are unrelated holds only if the 53 surveyed datasets are a fair and representative sample of the field, and the paper concedes some resources may have been missed.","fun_headline_variants_meta":{"raw":{"variants":["53 idiom datasets, zero bridges: type/token gap divides fields","Idiom studies: two silos, 53 datasets, no connection","Survey finds psycholinguistic and computational idiom data don't mix","Idiom datasets: psycholinguists and computationalists speak different languages","Type vs token: why idiom research has two incompatible worlds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000146,"raw_usage":{"total_tokens":962,"prompt_tokens":631,"completion_tokens":331,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":375,"completion_tokens_details":{"reasoning_tokens":240}},"tokens_in":375,"tokens_out":331,"duration_ms":3878,"temperature":1.0,"reasoning_tokens":240,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:41:47.875876+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Search the reference lists of the surveyed computational idiom papers for citations to psycholinguistic norming studies (or vice versa). Finding even one dataset built on both traditions — for example, a computational idiomaticity corpus whose labels were derived from psycholinguistic norming ratings, or a norming study that used a computational idiom corpus to select items — would weaken the paper's claim that the two fields have no relation.","supporting_citations":[],"review_version":1}