{"id":"cd0c2ca4-0861-4048-a4ed-a1f6ebf5cde0","arxiv_id":"2505.21301","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new Italian dataset of subordinate-level category exemplars shows that LLMs overlap with humans on only 10-24% of top available exemplars and that performance varies strongly across semantic domains.","lead":"Researchers collected 24,659 Italian words that people list as types of 187 everyday objects, then asked eight AI models to do the same task. They found that humans and AI agree on only about a quarter of the most typical examples, and agreement varies sharply by topic.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LLM 'most available' exemplars are ranked from only 5 generation runs, so the below-25% human overlap could be a small-sample artifact rather than a true categorical misalignment.","rationale":"Read in good faith, the paper is a solid, clearly reported descriptive study with a valuable new Italian dataset. The central claim is the low human-LLM alignment, quantified as 'below 25%' in §6. The reader's weakest assumption (the corpus-presence validity proxy in §4.1) is a real threat to the hallucination-rate analysis, but I see an even more direct threat to the alignment numbers: LLM availability is estimated from only 5 runs per category, whereas Eq. (4) was designed for dozens of human respondents. The resulting top-5 sets are noisy and tied, and no variance or chance analysis is reported. If this noise inflates the mismatch, the headline claim is overstated; if it does not, the claim is supported. This is eminently testable and does not require new data collection beyond additional model runs. I therefore keep the reader's conditional verdict: the paper should be accepted only after demonstrating that the overlap statistic is stable and above chance.","tokens_in":26197,"tokens_out":8849,"duration_ms":97820,"concrete_test":"For nemo-12B and llama-3.1-70B, run 50 generations per category on all 187 categories (or a stratified random sample of at least 60). Recompute availability via Eq. (4) and recompute Table 2 top-5 overlap for run counts 5, 10, 20, and 50 to test stabilization. Also build a chance baseline by sampling 5 items uniformly from each model's valid output vocabulary per category and computing the same overlap. If overlap stabilizes within 3 points of the reported values and the chance baseline is below 5%, the concern is resolved. If overlap rises substantially or the baseline approaches the reported overlap, the central 'low alignment' claim should be requantified or qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 2's overlap scores, which anchor the central 'below 25%' claim (§6), compare the 'most available' exemplars of humans and LLMs. Human availability comes from roughly 30 respondents per category (§3); LLM availability is obtained by plugging just 5 generation runs per category into Eq. (4) (Appendix B.5). With N=5, availability is quantized in units of 0.2 and contains many ties: any exemplar that appears first in at least one run receives the same score as others that appear first in exactly one run, regardless of how often it appears later; with 5 runs, up to 5 exemplars can tie at 0.2 by occupying the first position of different runs. The top-5 set is therefore partly arbitrary, and the paper reports no variance, stability, or chance baseline for the overlap percentages. The observed 10–24% overlaps could reflect sampling noise in the LLM ranking rather than genuine categorical misalignment; alternatively, even 24% could be near chance, making 'low alignment' uninformative. This is the most direct threat to the headline quantitative conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a new Italian psycholinguistic dataset of human-generated subordinate-level exemplars for 187 basic-level concrete categories, collected from 365 Italian speakers. It evaluates eight open LLMs and vLMs on exemplar generation, category induction, and typicality prediction. The headline result is a low alignment between human and LLM 'most available' exemplars, with top-5 overlap below 25% for all models, and higher validity for larger text models. The paper also reports that LLMs struggle with superordinate category induction and with typicality detection when categories have many similarly available exemplars.","tokens_in":26426,"tokens_out":7134,"duration_ms":71340,"significance":"If the results withstand the methodological concerns below, the paper would make a useful contribution: a novel psycholinguistic resource for Italian subordinate categories, an open comparison of multiple open models, and a clear demonstration that current open LLMs do not reproduce human subordinate-level availability. The data and code are promised to be released, which supports reproducibility. The closed-form subtasks (category induction, typicality) are a good complement to the generative task, and the qualitative hallucination taxonomy is insightful. However, the headline quantitative claim depends on a small-sample LLM availability estimate and lacks chance baselines, so the significance is conditional.","major_comments":[{"comment":"The central 'low alignment' result is based on LLM availability rankings computed from only five generation runs per category (Appendix B.5). With N=5, the availability score in Eq. (4) is quantized in units of 0.2 and contains many ties, so the top-5 most available set is partly arbitrary and unstable. The paper reports no variance, stability, or chance baseline for the overlaps, and the observed 10–24% values in Table 2 could reflect sampling noise in the LLM ranking rather than categorical misalignment. To support the headline claim, the authors should increase the number of runs, report bootstrap confidence intervals or stability across runs, and compare against a chance baseline such as random ranking of the model's valid exemplars.","section":"§4.2, Table 2; §6"},{"comment":"The operationalization of 'valid exemplar' as 'appears at least once in ItTenTen' conflates corpus attestation with semantic correctness. The paper itself shows that idefics2-8B lists other tree species (acacia, eucalyptus, maple) as exemplars of abete 'fir' (Section 4.2); these are corpus-attested but category-inappropriate. This conflation affects the interpretation of the validity percentages in Table 1, the hallucination taxonomy in Appendix B.8, and the composition of the 'valid' sets used in Table 2. The authors should validate a sample of generated items with human judgments and distinguish false negatives (rare but real expressions) from genuine hallucinations.","section":"§4.1, Table 1; §4.2"},{"comment":"Even if the sampling issue were fixed, the 'low alignment' claim would need a chance baseline to be meaningful. For categories with many possible subordinate labels, two independent rankings can have low top-5 overlap by chance; conversely, the human data themselves may have limited split-half reliability over the top-5 sets. The paper does not report any null-model comparison or human-human agreement, so the observed 24% could be close to chance. Adding such baselines is necessary to interpret the magnitude of the alignment.","section":"§6; Table 2"}],"minor_comments":[{"comment":"The paper uses 'basic-level categories' and 'subordinate level' inconsistently; the abstract says 'first attempt to examine the organization of categories by analyzing exemplars produced at the subordinate level,' while RQ1 and §2.2 say 'investigate the organization of basic-level categories.' Clarify that the stimuli are basic-level categories and the generated exemplars are subordinate-level labels.","section":"Abstract; §2.2; RQ1"},{"comment":"The overlap metric is not fully defined; specify the denominator (e.g., number of categories where at least one match occurs versus proportion of LLM top-n items appearing in human top-n) and the distinction between top-1 identity and set overlaps for top-3/top-5.","section":"Table 2"},{"comment":"The Italian prompt contains grammatical errors (e.g., 'elenca tutta i tipi di' should be 'elenca tutti i tipi di'; 'denota una concetto' should be 'denota un concetto'); correct the prompt if it is to be reused.","section":"Appendix B.5"},{"comment":"The Limitations section states 'Model are trained' (typo for 'Models') and Ethical Considerations uses 'Euro e 1.80' instead of '€1.80.'","section":"Limitations; Ethical Considerations"},{"comment":"The claim that 'all LLMs have not been exposed to these stimuli' rests on the unverifiable assumption that no training corpus contained similar data; this should be phrased as a plausibility argument rather than a guarantee.","section":"Ethical Considerations"},{"comment":"The sentence 'only 13.9% of the top-5 dominant exemplars overlap with the ranking of the top-5 most available exemplars' is ambiguous; it should specify whether this is the proportion of categories where the sets overlap exactly, overlap in any position, or the proportion of top-5 items that coincide.","section":"§3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the journal's scope and the dataset is a potentially reusable resource. However, the headline quantitative claim needs the additional analyses described in the major comments; without those, the paper's central conclusion is not yet supported. I would encourage the editors to request a revision addressing the sampling and baseline issues rather than rejecting, because the resource and the qualitative analyses are valuable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper's real contribution is the Italian dataset of subordinate-level exemplars for 187 basic-level categories, collected from 365 speakers. That is new and useful for psycholinguistic norming and for anyone benchmarking LLMs at a finer taxonomic grain than the usual superordinate level. The qualitative taxonomy of hallucination types (ad hoc, imitation-based, conceptual confusion) is also genuinely informative.\n\nThe central claim—low alignment between humans and LLMs—is directionally plausible and consistent with earlier work, but the headline numbers are weaker than the text suggests. Overlap is computed from just five generation runs per category. With N=5, the LLM availability score is quantized in steps of 0.2 and the top-5 set contains many ties, so the ranking is partly arbitrary. The paper gives no variance, stability, or chance baseline for the 10–24% overlap, and without a chance baseline \"low alignment\" is not a well-defined claim. That is the softest spot and it is load-bearing for Section 4.\n\nThe validity measure is also overclaimed. \"Valid\" means \"appears at least once in ItTenTen,\" not \"is a correct subordinate exemplar.\" The paper itself shows corpus-attested items that are category-inappropriate (e.g., idefics2 listing other tree species for 'fir'), so Table 1's percentages conflate lexical existence with semantic correctness. This is only partially acknowledged in the limitations and is not checked against human judgments.\n\nOn the other hand, the two subtasks (category induction and typicality judgment) are independent of the generation-run issue because they use human exemplars and perplexity scoring. Those results, including the finding that typicality prediction improves when the availability gap is larger, seem solid and are a useful addition.\n\nThe methods are clearly described, the literature is handled properly (Nighojkar, Heyman, Montefinese, Banks and Connell are all positioned correctly), and the inclusion of vLMs is a plus. The missing repository links are an editorial problem, not a scientific one.\n\nBottom line: this deserves peer review. A serious referee should ask for (a) chance baselines and variance estimates for the overlap scores, (b) a human check on a sample of corpus-attested exemplars, and (c) the actual data and code links. The dataset alone justifies moving forward.","headline":"Useful new Italian subordinate-level category dataset; the low-alignment claim is plausible but the overlap numbers need chance baselines and more than five LLM runs.","tokens_in":26941,"tokens_out":2495,"would_cite":true,"duration_ms":26381,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"For 187 basic-level Italian categories, open LLMs and humans produce different subordinate exemplars, with top-list overlap falling below 25%.","keywords":["subordinate categories","basic-level categories","exemplar generation","typicality","large language models","conceptual knowledge","Italian psycholinguistic norms","category induction"],"falsifier":"Take the 187 basic-level categories, ask a new sample of Italian speakers to rate each LLM-generated string (both corpus-attested and zero-frequency) as an acceptable, conventional type of X, and compare the resulting validity rates with the corpus-based ones; if speakers accept a substantial share of zero-frequency expressions, the reported hallucination rates are inflated. Alternatively, re-collect human exemplar norms from a second independent sample: if human-human top-5 overlap is as low as the model-human overlap (about 24%), the low-alignment result would reflect human production variability rather than an LLM-specific deficit.","tokens_in":1965,"feed_emoji":"🧠","tokens_out":2260,"duration_ms":101151,"temperature":0.7,"pith_summary":"This paper asks whether large language models organize conceptual knowledge the way people do one level down from everyday categories: when asked for kinds of 'dog' or 'chair,' people give subordinate exemplars like 'labrador,' 'German shepherd,' and 'dachshund.' The authors collected 24,659 such exemplars from 365 Italian speakers for 187 basic-level concrete words, producing what they describe as the first psycholinguistic norms for subordinate-level category organization in Italian. They then asked eight open language and vision-language models to generate exemplars for the same words and compared the results with human norms on availability, typicality, and category induction. The paper reports low alignment: the best model matches humans on about 24% of the five most available exemplars, and models frequently produce fluent but unattested ad hoc or implausible expressions. The upshot is a new resource plus evidence that current open LLMs are not reliable stand-ins for human conceptual structure at this level, while being useful for some categories.","feed_headline":"LLMs miss the subordinate categories Italians actually name","feed_subtitle":"New norms for 187 basic words show under 25% overlap on top exemplars, with models inventing ad hoc type names.","key_machinery":"The central machinery is the exemplar availability score, a measure drawn from established Italian semantic norms: for each basic-level category, availability weights how many participants produced an exemplar, at what position in their response list, and how early it appeared, yielding a graded gold-standard ordering of subordinate members. A second mechanism is corpus attestation via the ItTenTen web corpus: an LLM output counts as valid if it occurs at least once in that corpus, and zero-frequency strings are labelled hallucinations and qualitatively grouped into ad hoc, nonsensical, foreign-language, conceptual-confusion, and imitation-based patterns. A third mechanism is a perplexity-based forced-choice evaluation for category induction and typicality detection, in which models choose between candidate categories or typical/atypical sentences, making the comparison close-ended and reproducible. These three mechanisms together convert free-form lists into comparable ranked structures.","core_discovery":"The central discovery, on the paper's own terms, is that subordinate-level exemplar availability, how readily speakers produce a kind/type exemplar for a basic category, is a graded human structure that current open LLMs only partially capture. In generation, larger models produce up to 82% corpus-valid Italian expressions, but valid is not the same as correct: zero-frequency strings are often imitation-based extensions (red California fir), implausible composites (geranium with oak leaves), or ad hoc instances (hallway dresser). The overlap between human and machine top exemplars is below 25% even for the best model, and it is lowest for body parts and furnishing and highest for foods and animals. In forced-choice tasks, models identify the basic-level category of human exemplars almost perfectly (93-98% accuracy) but fail much more on the corresponding 12 superordinate categories (38-64%), and typicality detection drops as categories contain more similarly available exemplars. Vision-language models do not close the gap: one vision model even lists other basic-level tree species as subordinate exemplars of fir.","pith_inferences":["The corpus-attestation filter likely cuts both ways: zero-frequency strings may include perfectly acceptable novel Italian compounds, while corpus-attested strings can be category-inappropriate (a different tree species listed under fir), so both the validity rates and the hallucination counts probably bracket true semantic correctness rather than measure it; a human acceptability rating study on ","The human and machine generation tasks are not perfectly matched: humans free-list in a self-paced survey, while models receive a few-shot prompt with an appliance example, so part of the low top-5 overlap could come from instruction differences rather than from different category stores; matching the elicitation format more closely would isolate the structural claim.","If the paper's domain-level pattern holds, a testable prediction follows: for categories dominated by encyclopedic or textual knowledge (foods, vehicles), future model generations will approach human overlap, while for categories grounded in bodily experience (body parts) or object geometry (furnishing), the gap will persist.","The results imply that valid exemplar is not a single concept: a string can be well-formed, attested, intended, and correct, and these four dimensions come apart in LLM output; future datasets could score each dimension separately."],"forward_implications":["The new Italian dataset extends existing Italian semantic norms with subordinate-level exemplars plus dominance, availability, mean rank order, and first-occurrence measures for 187 basic-level concepts.","Because valid-but-not-correct exemplars are frequent, NLP pipelines that use LLM-generated subordinate terms (ontology population, knowledge-base construction, vocabulary teaching) will need human-in-the-loop verification.","The five hallucination patterns provide a checklist for evaluating category-aware generation: ad hoc instances, nonsensical composites, foreign-language items, conceptual confusion, and imitation-based overgeneralization.","The sharp drop from basic-level to superordinate-level category induction (93-98% down to 38-64%) implies that more general taxonomic reasoning is a specific weakness of these models, not a general categorization failure.","Typicality judgments are recoverable when the availability gap between typical and atypical exemplars is large, but they flatten as categories accumulate many similarly available members, so LLMs mirror human typicality only where human structure is sharply graded."],"supporting_citations":[{"why":"Provides the 187 basic-level categories and the dominance, availability, mean rank, and first-occurrence measures that the new human dataset extends.","marker":"Montefinese et al. 2012"},{"why":"Establishes the basic-level/subordinate taxonomy and the shared-attribute claim that motivates studying subordinate categories.","marker":"Rosch et al. 1976"},{"why":"Defines the ItTenTen corpus used to separate valid exemplars from zero-frequency hallucinations.","marker":"Jakubíček et al. 2013"},{"why":"Documents the web-crawling method behind ItTenTen, the corpus used as the validity filter.","marker":"Suchomel et al. 2012"},{"why":"Prior transformer-based semantic fluency baseline that the generation task and low-overlap results build on.","marker":"Nighojkar et al. 2022"},{"why":"Shows ChatGPT matches human typicality ratings, the benchmark that the typicality subtasks test against.","marker":"Heyman and Heyman 2024"},{"why":"Shows textual models beat vision models on human typicality, the comparison that the vision-language model results extend.","marker":"Vemuri et al. 2024"},{"why":"Defines ad hoc categories, one of the hallucination patterns used in the qualitative analysis.","marker":"Barsalou 1983"}],"fun_headline_variants":["LLM subordinate categories: under 25% overlap with humans","Vision models can't fix LLM subordinate category gap","AI invents ad hoc subordinates, missing Italian norms","Subordinate category test: LLMs score low on Italian data","LLMs fail to capture human subordinate category structure"],"cache_read_input_tokens":29184,"weakest_assumption_plain":"The load-bearing premise is that a generated expression counts as a valid Italian subordinate term exactly when it appears at least once in the ItTenTen web corpus, with corpus absence marking hallucination; this proxy is never checked against human judgments, and the paper itself shows corpus-attested items can be category-inappropriate.","fun_headline_variants_meta":{"raw":{"variants":["LLM subordinate categories: under 25% overlap with humans","Vision models can't fix LLM subordinate category gap","AI invents ad hoc subordinates, missing Italian norms","Subordinate category test: LLMs score low on Italian data","LLMs fail to capture human subordinate category structure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00019,"raw_usage":{"total_tokens":1324,"prompt_tokens":915,"completion_tokens":409,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":330}},"tokens_in":531,"tokens_out":409,"duration_ms":5956,"temperature":1.0,"reasoning_tokens":330,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:30:58.970637+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the 187 basic-level categories, ask a new sample of Italian speakers to rate each LLM-generated string (both corpus-attested and zero-frequency) as an acceptable, conventional type of X, and compare the resulting validity rates with the corpus-based ones; if speakers accept a substantial share of zero-frequency expressions, the reported hallucination rates are inflated. Alternatively, re-collect human exemplar norms from a second independent sample: if human-human top-5 overlap is as low as the model-human overlap (about 24%), the low-alignment result would reflect human production variability rather than an LLM-specific deficit.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the 187 basic-level categories and the dominance, availability, mean rank, and first-occurrence measures that the new human dataset extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the basic-level/subordinate taxonomy and the shared-attribute claim that motivates studying subordinate categories."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the web-crawling method behind ItTenTen, the corpus used as the validity filter."},{"cited_title":"Cognitive Modeling of Semantic Fluency Using Transformers","cited_arxiv_id":"2208.09719","evidence_quote":"Prior transformer-based semantic fluency baseline that the generation task and low-overlap results build on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows ChatGPT matches human typicality ratings, the benchmark that the typicality subtasks test against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows textual models beat vision models on human typicality, the comparison that the vision-language model results extend."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines ad hoc categories, one of the hallucination patterns used in the qualitative analysis."}],"review_version":1}