{"id":"2a005be4-f49e-402f-b013-8dba7d55197c","arxiv_id":"2607.17221","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Cosmetics, Crayola, and car color vocabularies occupy the shared 86-category COLIBRI color space differently, with Crayola broadest and most balanced, cosmetics biased toward warm tones, and cars toward blue and achromatic tones.","lead":"This paper maps color names from cosmetics, crayons, and car paints onto a shared fuzzy color space and shows that each domain uses different regions of that space. It offers a simple quantitative way to build context-aware color vocabularies for product search, design, and AI systems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim that color naming cannot be explained by numerical color similarity lacks any similarity baseline or statistical test.","rationale":"The reader's weakest assumption focuses on dataset curation and sample-size imbalance, which are legitimate threats to the robustness of the specific rankings (Crayola broadest, Car narrowest). However, the single most load-bearing concern is more fundamental: even if the datasets are perfectly representative, the paper's experimental design never tests the central interpretive claim. The paper shows that domain-specific vocabularies occupy the COLIBRI space differently, but this is an observation about the vocabularies themselves, not about human color naming relative to numerical similarity. To conclude that 'color naming cannot be fully explained by numerical color similarity alone,' one must demonstrate that numerical similarity is insufficient in a predictive or counterfactual sense—e.g., that colors closer than some threshold receive different names across contexts, or that a similarity-only model makes incorrect predictions that a context-aware model corrects. No such demonstration appears. The concern is addressable (add a baseline and permutation tests), so a conditional verdict remains appropriate. I therefore leave the reader's verdict unchanged, while noting that the justification for conditionality should be broadened to include this missing baseline, not only dataset curation issues.","tokens_in":7053,"tokens_out":3839,"duration_ms":40458,"concrete_test":"Construct a numerical-similarity null model for each domain: map every RGB sample to the nearest COLIBRI category center in CIELAB ΔE, and compute the resulting coverage, normalized entropy, and max lift under the same name-assignment counts. Compare observed metrics to the null distribution via bootstrap (resampling names within each domain). Additionally, identify cross-domain RGB pairs with ΔE < 2 and test whether their associated names fall into different COLIBRI categories more often than expected by chance. If the null model reproduces the observed differences, the central claim fails; if not, the context-specific effect is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central conclusion is that 'color naming cannot be fully explained by numerical color similarity alone.' Yet the analysis in Section III-B and Section IV only maps three curated vocabularies onto the fixed COLIBRI partition and reports descriptive statistics (coverage, normalized entropy, max lift). There is no comparison against any numerical-similarity baseline, no test of whether physically close colors receive different names across domains, and no statistical significance testing. Because the vocabularies were sourced from different domains, it is trivially expected that their name distributions differ; the observed differences could arise entirely from the source of the vocabulary while numerical similarity still perfectly determines the RGB-to-category mapping. Thus the load-bearing claim that semantic context matters 'beyond numerical similarity' is internally under-supported by the reported evidence, independent of the dataset curation concerns raised by the reader.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether semantic context changes how color vocabularies occupy a shared perceptual color space. It maps three curated color-name datasets—cosmetics shades, Crayola names, and car paint names—onto the 86 fuzzy color categories of the COLIBRI model via an RGB→HSI→COLIBRI pipeline. For each context it reports coverage, normalized Shannon entropy, and maximum lift, and assigns a qualitative richness class (e.g., broad-balanced, narrow-specialized). The main reported findings are that Crayola has the broadest and most balanced coverage, cosmetics is concentrated in warm regions, and car colors are concentrated in blue and achromatic regions. The abstract and conclusion further claim that color naming cannot be fully explained by numerical color similarity alone and that semantic context plays a central role.","tokens_in":7316,"tokens_out":3724,"duration_ms":39272,"significance":"If the central claim were established, the paper would make a modest but useful contribution: a simple, transparent descriptive framework for comparing domain-specific color vocabularies on a fixed fuzzy-color representation, with potential applications in design analytics and context-aware color naming. The metrics used are elementary and easy to compute, and the paper is candid about some limitations. However, the strongest claimed conclusion—that semantic context matters beyond numerical color similarity—is not actually tested by the reported analyses. The manuscript also does not ship data, code, or confidence intervals, and the dataset curation is subjective. These issues are fixable, and the paper's descriptive statistics, if accompanied by appropriate baselines and uncertainty quantification, could form a sound empirical contribution.","major_comments":[{"comment":"The load-bearing conclusion that 'color naming cannot be fully explained by numerical color similarity alone' is not supported by the analysis. Sections III-B and IV only map each dataset onto the shared COLIBRI partition and compute aggregate distributions (coverage, entropy, lift). There is no baseline that uses numerical color similarity alone to predict or explain the naming patterns, no test of whether physically close colors receive different names across contexts, and no statistical comparison. The observed differences could arise entirely from the source corpora while the RGB-to-category mapping remains fully determined by numerical color similarity. I recommend adding a concrete similarity-based baseline—e.g., predicting assigned COLIBRI categories from color coordinates alone, or showing that the same COLIBRI region/Lab neighborhood receives significantly different context-spec","section":"Abstract and Section IV"},{"comment":"The richness ranking (Crayola broadest, Car narrowest) is not robust because the datasets have very different sizes—Car has 297 records, Cosmetics 176, Crayola 169—and no confidence intervals, bootstrap, or rarefaction are provided. Coverage and entropy are sample-size dependent, so the ranking could be an artifact. The paper's own limitation paragraph compounds this by saying that Car's lower coverage may reflect its 'smaller name pool,' but Table III shows Car has the largest record count; this is internally inconsistent. Add rarefaction curves, per-category confidence intervals, or a matched-subsample comparison before using these rankings as evidence.","section":"Table III and Section IV"},{"comment":"The dataset curation is a load-bearing part of the analysis. The paper excludes shade names that are 'primarily evocative or marketing-driven rather than descriptive' but gives no operational criteria, no list of excluded names, and does not name the cosmetics retailer or provide a dataset link. Since the core comparisons concern vocabulary composition, this subjective filter could create or exaggerate the observed concentration in cosmetics warm tones. I recommend releasing the raw and filtered datasets and performing a sensitivity analysis that includes all names, or at minimum documenting the exclusion rule sufficiently for replication.","section":"Section III-A"},{"comment":"The 'color naming richness classification' is stated as a contribution, but the thresholds for 'high/moderate/low' coverage, normalized entropy, and maximum lift are never specified. The rule-based labels such as broad-concentrated and narrow-specialized are therefore not reproducible and do not add information beyond the raw metrics. Define the cutoff values explicitly, or replace the labels with a transparent scoring rule.","section":"Section III-B2 and Table IV"}],"minor_comments":[{"comment":"The last paragraph of the introduction says the paper reports 'experimental results across the four semantic contexts,' but only three contexts (Cosmetics, Crayola, Car) are analyzed. Correct to 'three.'","section":"Section I"},{"comment":"Equation (2) has a dangling 'Nd =' in the typeset text; the formatting should be cleaned up. Also, the definitions of pk,d and the effective number of fuzzy colors could be stated more clearly.","section":"Equations"},{"comment":"Figures 1–3 are referenced only loosely in the text; the pipeline in Fig. 1 and the distribution plots in Fig. 3 deserve at least a sentence of explicit interpretation. Figure 3's y-axis says 'number of distinct color names,' which should be defined consistently with the 'Total Count' column in Table III.","section":"Figures"},{"comment":"References [9] and [21] are the same work (arXiv and ICPR versions); consolidate them to avoid duplication. Also, the paper does not provide a data/code availability statement; if the dataset can be shared, add one.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The COLIBRI model [7] is authored by the same research group and used as the fixed ground-truth perceptual space. This is not circularity in the main derivation, but it means the central measurement scale is not independently verified in this manuscript. Independent replication or at least a clear description of COLIBRI's membership functions would strengthen confidence. Additionally, the central 'beyond numerical similarity' claim appears to require a baseline comparison that is absent; the revision should prioritize that over stylistic improvements."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a small, internally consistent descriptive study whose headline claim is not supported by its design. The useful part is a reusable way to compare domain color vocabularies on a shared fuzzy space; the overreach is the assertion that color naming cannot be explained by numerical similarity alone.\n\nWhat's actually new: the specific head-to-head comparison of Cosmetics, Crayola, and Car color names on the 86-category COLIBRI space—coverage, normalized entropy, max lift—is not in the cited literature. The metrics are standard, but they are applied cleanly and the richness classification is transparent. I see no arithmetic errors. The paper also admits sample-size sensitivity, which is more than most short papers do.\n\nWhere it gets soft: the central claim. Nothing in the paper compares against a numerical-similarity baseline. The authors show that three vocabularies occupy COLIBRI differently. That is a statement about the vocabularies, not about whether numerical similarity can explain naming. If the same RGB maps to different names in different domains, that would be evidence; they do not test that. So the 'cannot be fully explained' sentence is unsupported. This is a real flaw, not a nitpick.\n\nThe data is not shipped, sample sizes are unequal (176/169/297), no confidence intervals or rarefaction are provided, and the exclusion of 'evocative' shade names is subjective. Those are fixable. The COLIBRI self-citation is fine—it is a measurement scale, not a circular derivation.\n\nWho this is for: people working on color naming, visualization, or product search who want a quick, cheap way to compare domain vocabularies. It will not change anyone's theory, but the framework is reusable.\n\nRecommendation: send it to peer review with a clear expectation of revision. The authors need to either soften the claim to 'the three vocabularies differ' or add an actual similarity-based baseline and statistical testing. With data and a more modest conclusion, this could become a solid short paper.","headline":"Small, internally consistent descriptive study whose headline claim outruns its design; worth refereeing for the reusable framework, not for the conclusion as stated.","tokens_in":7747,"tokens_out":2357,"would_cite":false,"duration_ms":24215,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper demonstrates that semantic context changes how a color vocabulary occupies perceptual color space, and quantifies the difference by mapping cosmetics, Crayola, and car-color names onto the 86 fuzzy categories of the COLIBRI model","keywords":["color naming","semantic context","fuzzy color model","COLIBRI","Shannon entropy","coverage","lift","domain-specific vocabulary"],"falsifier":"Equalize sample sizes by rarefaction or include the excluded evocative cosmetics names and recompute coverage: if the Crayola-broadest / car-narrowest ordering disappears, or if a random sample of generic color names reproduces the same coverage pattern, the central claim would be undercut.","tokens_in":7019,"feed_emoji":"🎨","tokens_out":4053,"duration_ms":37006,"temperature":0.7,"pith_summary":"The paper sets out to show that the words people use for colors depend on the semantic domain they are talking about, and that this difference can be measured in one shared perceptual color space. It maps three real-world color vocabularies—cosmetics shades, Crayola crayon names, and car paint names—into the 86 fuzzy color categories of the COLIBRI model, then compares them by coverage, entropy, and lift. The three vocabularies occupy the space differently: Crayola spreads across the most categories with the most even distribution, cosmetics clusters in warm reds and oranges, and car paints concentrate in blues and achromatic neutrals. The paper argues that color naming therefore cannot be predicted from numerical color similarity alone, and that context-aware color vocabularies should replace a single universal color-name set.","feed_headline":"Crayola covers 50 of 86 color regions; car paints just 40","feed_subtitle":"Cosmetics cluster in warm reds, cars in blues and grays—color naming can't rely on RGB distance alone.","key_machinery":"The carrying mechanism is the COLIBRI fuzzy color model, which partitions perceptual color space into 86 soft, overlapping color categories rather than hard bins; each color sample is converted via RGB→HSI→COLIBRI and assigned soft membership. On top of this, the paper defines a rule-based color naming richness classification using three indicators: coverage (fraction of the 86 regions touched), normalized Shannon entropy (evenness of name distribution), and maximum lift (how strongly a single region is overrepresented relative to uniform use). Together these turn 'how does a domain talk about color' into three comparable numbers.","core_discovery":"On the paper's own terms, the discovery is that semantic context leaves a measurable fingerprint in perceptual color space. After mapping RGB samples to HSI and then to COLIBRI's 86 fuzzy categories, Cosmetics covers 48 categories, Crayola covers 50, and Car colors cover 40. Crayola has the highest normalized entropy (0.83) and effective fuzzy colors (39.66), car colors the lowest (0.70 and 22.29), and cosmetics shows a maximum lift of 14.17 on red-orange regions while car colors show lifts concentrated on blue and gray regions. The same fixed color space thus hosts different domain vocabularies with different breadth, balance, and focus.","pith_inferences":["A testable extension: run human color-naming experiments in each domain and check whether people's agreement patterns match the fuzzy-category distributions reported here; if they do, the framework becomes a predictor of naming behavior, not just a descriptive tool.","If the pattern holds across more domains, the shape of a domain's color vocabulary—broad versus narrow, warm versus cool—might be predictable from the typical colors of objects in that domain and their marketing function, extending information-theoretic accounts of color naming.","The framework could be inverted for applications: given a product image, infer the likely domain vocabulary and use it to generate candidate color names for search or recommendation, which the paper hints at but does not implement."],"forward_implications":["Color similarity systems that ignore semantic context will mispredict how colors are named in real-world domains.","Domain-specific color vocabularies are more appropriate than a single universal color-name set for product search, recommendation, and design analytics.","The same fuzzy color representation can host many naming granularities, so a system can keep a fixed perceptual base and vary only the semantic vocabulary.","The coverage-entropy-lift scheme can classify any new color vocabulary as broad-balanced, broad-concentrated, narrow-specialized, or sparse-low richness."],"fun_headline_variants":["Crayola spans 50 color regions, car paints just 40","Color naming depends on domain: Crayola broadest, cars specialized","Semantic context shapes color words more than RGB","From Crayola to car paint: context changes color names","Cosmetics warm, cars cool: context drives color terms"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The ranking depends on the curated datasets being representative of each domain: shade names that were 'primarily evocative or marketing-driven' were removed, and the three datasets differ in size (176, 169, and 297 records) with no rarefaction, so the observed coverage differences could partly reflect curation and sample size rather than true domain structure.","fun_headline_variants_meta":{"raw":{"variants":["Crayola spans 50 color regions, car paints just 40","Color naming depends on domain: Crayola broadest, cars specialized","Semantic context shapes color words more than RGB","From Crayola to car paint: context changes color names","Cosmetics warm, cars cool: context drives color terms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000563,"raw_usage":{"total_tokens":2501,"prompt_tokens":731,"completion_tokens":1770,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":475,"completion_tokens_details":{"reasoning_tokens":1684}},"tokens_in":475,"tokens_out":1770,"duration_ms":11568,"temperature":1.0,"reasoning_tokens":1684,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T18:38:23.556927+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Equalize sample sizes by rarefaction or include the excluded evocative cosmetics names and recompute coverage: if the Crayola-broadest / car-narrowest ordering disappears, or if a random sample of generic color names reproduces the same coverage pattern, the central claim would be undercut.","supporting_citations":[],"review_version":1}