{"id":"9ded2ed2-c932-44b1-9f23-ee56de392d86","arxiv_id":"2607.10465","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In free multilingual naming of 18 COLIBRI shades, green is named consistently far more often than red or yellow, indicating unequal perceptual category stability.","lead":"Green color names stay stable across many shades while yellow names scatter into brown, gold, and mustard in a free-naming study with 92 speakers. The ranking of category robustness can guide better color models for interfaces and vision systems.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"Consistency ranking may be an artifact of the hand-built multilingual mapping rather than of perceptual category breadth.","rationale":"The reader correctly isolates the hand-crafted dictionary (Table II / Algorithm 1) as the weakest assumption. The free-naming design is appropriate and the raw word clouds already suggest greater lexical diversity for yellow, so the directional claim is not baseless; yet the quantitative ranking that constitutes the strongest claim is produced solely by that unvalidated mapping. No other methodological issue (convenience sample, uncontrolled displays, lack of inferential tests) is as directly load-bearing for the numerical order itself. A blinded re-mapping is a cheap, decisive check. Until it is performed, the CONDITIONAL verdict remains appropriate; the concern does not justify rejection because the qualitative pattern is still visible in the unprocessed responses.","tokens_in":9480,"tokens_out":479,"duration_ms":7449,"concrete_test":"Have two independent bilingual annotators (blind to the paper’s ranking) re-map every free response using only the raw strings and a minimal instruction set that does not privilege the authors’ Table II. Recompute the three consistency scores. If the Green > Red > Yellow order reverses or the gaps shrink below 5 percentage points under either annotator’s mapping, the ranking is mapping-dependent and the strongest claim weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is the ordered consistency Green (65.55%) > Red (45.87%) > Yellow (39.06%), interpreted as evidence that green occupies a broader perceptual region. That percentage is produced by Algorithm 1 steps 5–6: free responses are first normalized, then forced into a discrete category via the hand-crafted lexicon of Table II. The lexicon is not validated against independent annotators or against a held-out naming corpus; many borderline terms (salad/lime/swamp for green; mustard/beige/golden/brown for yellow; pink/coral/burgundy for red) are assigned by author judgment. Because yellow receives a larger share of such borderline terms, any systematic over-assignment of those terms away from the target hue will mechanically lower yellow’s consistency relative to green. Without an inter-annotator reliability check or an alternative mapping, the reported ranking cannot be distinguished from a mapping artifact.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper reports a free color-naming experiment (n=92; Kazakh, Russian, English) on 18 COLIBRI-derived stimuli spanning red, yellow, and green at three saturations and two intensities. Free responses are normalized, language-detected, and mapped via an author-built lexicon (Algorithm 1, Table II) to discrete categories; consistency is the fraction of responses that map back to the stimulus hue. The central empirical claim is an ordered naming consistency Green (65.55%) > Red (45.87%) > Yellow (39.06%), with supporting word clouds, heatmaps, and saturation/intensity breakdowns, interpreted as evidence that green occupies a broader, more robust perceptual region than yellow.","tokens_in":9707,"tokens_out":823,"duration_ms":22408,"significance":"If the Green > Red > Yellow stability ordering is robust to mapping choices and viewing conditions, the result is useful for perceptually grounded color models, multilingual color-naming systems, and HCI design that must tolerate appearance variation. Strengths include free (not forced-choice) naming, multilingual coverage, transparent frequency tables and heatmaps, and explicit reporting of object-based descriptors. The contribution is empirical and applied rather than theoretical; its value depends on showing that the ranking is not an artifact of the hand-built lexicon or uncontrolled displays.","major_comments":[{"comment":"Algorithm 1 (steps 5–6) and Table II: consistency Ch is defined as the fraction of free responses that map to the COLIBRI hue via a hand-crafted multilingual lexicon. Many borderline terms (salad/lime/swamp; mustard/beige/golden/brown; pink/coral/burgundy) are assigned by author judgment with no inter-annotator reliability, alternative mapping, or sensitivity analysis. Because yellow attracts more such borderline terms, systematic over-assignment away from yellow would mechanically produce the reported Green > Red > Yellow order. The central ranking cannot be distinguished from a mapping artifact without at least (i) dual independent annotation of a response sample and (ii) a leave-one-mapping-rule-out or coarser/finer lexicon reanalysis.","section":null},{"comment":"Table III and §IV: the headline percentages (65.55%, 45.87%, 39.06%) are reported without confidence intervals, bootstrap uncertainty, or any inferential test of the Green > Red > Yellow ordering (or of saturation/intensity effects in Fig. 4). With ~320 responses per hue, such tests are feasible; without them the claim that categories “differ in their consistency” remains descriptive only and is not yet load-bearing evidence for unequal category breadth.","section":null},{"comment":"§III Methodology: stimuli were shown in an online form with no display calibration, white-point control, or lighting instructions. For a color-perception claim about saturation/intensity robustness, uncontrolled sRGB rendering and mixed devices can systematically shift low-saturation and medium-intensity samples (especially yellows) toward brown/gray/beige—the very alternatives that lower yellow consistency. Either restrict claims to “naming under typical web viewing” or add a controlled lab/replication subset; as written, the perceptual-stability interpretation overreaches the uncontrolled setup.","section":null},{"comment":"Table I and stimulus selection: all 18 samples come from the authors’ own COLIBRI fuzzy model (self-cited), and consistency is recovery of the COLIBRI hue label. Free responses supply independent data, so this is not pure circularity, but the grid may over-sample regions where COLIBRI already treats green as broad and yellow as narrow. A brief comparison to an independent atlas (e.g., Munsell or WCS-style chips) or an explicit statement that results are COLIBRI-conditioned would clarify external validity of the “broader perceptual region” claim.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The new piece here is a clean free-naming comparison of naming stability for red, yellow, and green across a controlled saturation × intensity grid, with Kazakh responses included alongside Russian and English. That is a legitimate, usable extension of the Berlin–Kay / Regier / Witzel / Mylonas line rather than a new theory.\n\nWhat they did well: 92 participants, free responses (not forced choice), transparent preprocessing, word clouds, heatmaps, and a clear consistency order Green 65.55% > Red 45.87% > Yellow 39.06%. The saturation effect (high 73.65% vs low 25.24%) is intuitive and well illustrated. The semantic breakdown (food/plant/earth associations) is a nice secondary observation. Citations cover the right literature and the COLIBRI self-citation is appropriate given that the stimuli come from it.\n\nSoft spots, in proportion. The load-bearing step is Algorithm 1 + Table II: free multilingual strings are collapsed by a hand-built lexicon into discrete categories, and consistency is simply the fraction that land back on the COLIBRI hue. Borderline terms (salad/lime/swamp; mustard/beige/golden/brown; pink/coral/burgundy) are assigned by author judgment with no inter-annotator check or alternative mapping. That is exactly the stress-test concern, and it is real: if yellow attracts more of those borderline labels and they are systematically mapped away from yellow, the ranking is partly artifactual. Convenience sample (conference attendees), no display calibration, and no inferential statistics or CIs are secondary but real limitations. Circularity is mild because free responses are independent data; the mapping is the weaker link.\n\nWho it is for: people building fuzzy color models or naming systems who need empirical stability numbers, especially with a Central Asian language sample. Not for theorists looking for a new mechanism.\n\nI would send it to referees. The design is appropriate, the data are new, and the limitations are fixable (release the raw responses, report reliability on the mapping, add simple stats). Worth engaging if you work on color naming or perceptual color models; not essential otherwise.","headline":"Solid free-naming data showing Green > Red > Yellow consistency under saturation/intensity variation; the ranking is useful but rests on an unvalidated author dictionary and a convenience sample.","tokens_in":10298,"tokens_out":540,"would_cite":false,"duration_ms":8144,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Green stays green under shade changes; yellow does not, and red sits in between.","keywords":["color perception","color naming","color categorization","categorical perception","cross-linguistic variation","perceptual color space","naming consistency","COLIBRI"],"falsifier":"Re-run the same free-naming design on a larger, balanced sample with an independently validated mapping (or no forced binning), and check whether the Green > Red > Yellow consistency order still holds for matched saturation and intensity levels.","tokens_in":10409,"feed_emoji":"🎨","tokens_out":802,"duration_ms":12936,"temperature":0.7,"pith_summary":"This paper asks how much a color can change in saturation and brightness and still keep the same name. In a free naming task with 92 speakers of Kazakh, Russian, and English, participants named 18 red, yellow, and green samples taken from a perception-based color model. Naming consistency ranked Green (about two-thirds of answers stayed \"green\") above Red (under half) and Yellow (under two-fifths), with yellow often sliding into brown, gold, mustard, or beige. The authors argue that some color categories occupy broader regions of perceptual space and are therefore more stable under visual variation, which matters for anyone building color naming systems or models meant to match human perception.","feed_headline":"Green holds its name; yellow frays under shade changes","feed_subtitle":"Free naming by 92 speakers ranks category stability: green, then red, then yellow.","key_machinery":"Consistency score Ch: after free responses are normalized and mapped through a multilingual color lexicon, Ch is the fraction of answers for a hue that land in that hue's own category. The score, plus heatmaps and word clouds by hue, saturation, and intensity, is what ranks the categories.","core_discovery":"Color categories are not equally stable under changes in appearance. Across free multilingual names for 18 red, yellow, and green shades, consistency followed Green (65.55%) > Red (45.87%) > Yellow (39.06%). Green remained largely \"green\" despite shade differences; yellow produced many alternative labels including brown-, gold-, and mustard-related terms; red fell in between and often drifted toward pink or coral when desaturated or lightened.","pith_inferences":["If green really occupies a wider region, compression or gamut-mapping algorithms that protect green prototypes may preserve nameability better than equal treatment of all primaries.","The same ranking, if it generalizes, would predict higher cross-language agreement for green than for yellow in other free-naming corpora.","Borderline yellow-to-brown and red-to-pink mappings are natural places to test whether soft (fuzzy) category membership improves automated color naming over hard bins."],"forward_implications":["Perceptually grounded color models should treat green as a broader, more forgiving category than yellow.","Color naming systems will get higher agreement on green shades and more label diversity on yellow shades under the same appearance changes.","Saturation and intensity matter differently by hue: high saturation stabilizes naming most, while yellow is especially fragile at medium intensity.","Object-based and food-based descriptors are a large share of free names and should be expected in real-world color interfaces."],"fun_headline_variants":["Green stays green; yellow splinters into gold brown mustard","Color names: green most stable red mid yellow least consistent","Free naming ranks green over red over yellow for shade stability","Yellow frays most as shades shift; green holds across 18 samples","Green 65% red 46% yellow 39%: free multilingual color naming"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The hand-built dictionary that turns free multilingual answers—including food, plant, and shade phrases—into a few color bins must recover true category membership without systematically pushing yellow or red off their labels.","fun_headline_variants_meta":{"raw":{"variants":["Green stays green; yellow splinters into gold brown mustard","Color names: green most stable red mid yellow least consistent","Free naming ranks green over red over yellow for shade stability","Yellow frays most as shades shift; green holds across 18 samples","Green 65% red 46% yellow 39%: free multilingual color naming"]},"model":"grok-4.5","effort":"low","cost_usd":0.004264,"raw_usage":{"total_tokens":1234,"prompt_tokens":785,"num_sources_used":0,"completion_tokens":91,"cost_in_usd_ticks":42640000,"prompt_tokens_details":{"text_tokens":785,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":358,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":785,"tokens_out":91,"duration_ms":6032,"temperature":1.0,"reasoning_tokens":358,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T11:30:22.284299+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same free-naming design on a larger, balanced sample with an independently validated mapping (or no forced binning), and check whether the Green > Red > Yellow consistency order still holds for matched saturation and intensity levels.","supporting_citations":[],"review_version":1}