Pith. sign in

REVIEW 4 major objections 31 references

Not All Color Categories Are Equally Stable: A Multilingual Free Color Naming Experiment

T0 review · 4 major / 0 minor · reviewed 2026-07-14 · grok-4.5

Pith's one-line read Green stays green under shade changes; yellow does not, and red sits in between.

desk verdict Solid free-naming data showing Green > Red > Yellow consistency under saturation/intensity variation; the ranking is useful but rests on an unvalidated author dictionary and a convenience sample. read the letter →

arxiv 2607.10465 v1 pith:MK26LGF3 submitted 2026-07-11 cs.CV cs.CL

classification cs.CVcs.CL
keywords colorperceptionnamingcategorizationcategoricalcross-linguisticvariationperceptualspaceconsistencyCOLIBRI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks how much a color can change in saturation and brightness and still keep the same name. In a free naming task with 92 speakers of Kazakh, Russian, and English, participants named 18 red, yellow, and green samples taken from a perception-based color model. Naming consistency ranked Green (about two-thirds of answers stayed "green") above Red (under half) and Yellow (under two-fifths), with yellow often sliding into brown, gold, mustard, or beige. The authors argue that some color categories occupy broader regions of perceptual space and are therefore more stable under visual variation, which matters for anyone building color naming systems or models meant to match human perception.

What carries the argument

Consistency score Ch: after free responses are normalized and mapped through a multilingual color lexicon, Ch is the fraction of answers for a hue that land in that hue's own category. The score, plus heatmaps and word clouds by hue, saturation, and intensity, is what ranks the categories.

What would settle it

Re-run the same free-naming design on a larger, balanced sample with an independently validated mapping (or no forced binning), and check whether the Green > Red > Yellow consistency order still holds for matched saturation and intensity levels.

Watch

Extended reading notes

Core claim

Color categories are not equally stable under changes in appearance. Across free multilingual names for 18 red, yellow, and green shades, consistency followed Green (65.55%) > Red (45.87%) > Yellow (39.06%). Green remained largely "green" despite shade differences; yellow produced many alternative labels including brown-, gold-, and mustard-related terms; red fell in between and often drifted toward pink or coral when desaturated or lightened.

Load-bearing premise

The hand-built dictionary that turns free multilingual answers—including food, plant, and shade phrases—into a few color bins must recover true category membership without systematically pushing yellow or red off their labels.

Editorial extensions

If this is right

  • Perceptually grounded color models should treat green as a broader, more forgiving category than yellow.
  • Color naming systems will get higher agreement on green shades and more label diversity on yellow shades under the same appearance changes.
  • Saturation and intensity matter differently by hue: high saturation stabilizes naming most, while yellow is especially fragile at medium intensity.
  • Object-based and food-based descriptors are a large share of free names and should be expected in real-world color interfaces.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If green really occupies a wider region, compression or gamut-mapping algorithms that protect green prototypes may preserve nameability better than equal treatment of all primaries.
  • The same ranking, if it generalizes, would predict higher cross-language agreement for green than for yellow in other free-naming corpora.
  • Borderline yellow-to-brown and red-to-pink mappings are natural places to test whether soft (fuzzy) category membership improves automated color naming over hard bins.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 0 minor

Summary. The paper reports a free color-naming experiment (n=92; Kazakh, Russian, English) on 18 COLIBRI-derived stimuli spanning red, yellow, and green at three saturations and two intensities. Free responses are normalized, language-detected, and mapped via an author-built lexicon (Algorithm 1, Table II) to discrete categories; consistency is the fraction of responses that map back to the stimulus hue. The central empirical claim is an ordered naming consistency Green (65.55%) > Red (45.87%) > Yellow (39.06%), with supporting word clouds, heatmaps, and saturation/intensity breakdowns, interpreted as evidence that green occupies a broader, more robust perceptual region than yellow.

Significance. If the Green > Red > Yellow stability ordering is robust to mapping choices and viewing conditions, the result is useful for perceptually grounded color models, multilingual color-naming systems, and HCI design that must tolerate appearance variation. Strengths include free (not forced-choice) naming, multilingual coverage, transparent frequency tables and heatmaps, and explicit reporting of object-based descriptors. The contribution is empirical and applied rather than theoretical; its value depends on showing that the ranking is not an artifact of the hand-built lexicon or uncontrolled displays.

major comments (4)
  1. Algorithm 1 (steps 5–6) and Table II: consistency Ch is defined as the fraction of free responses that map to the COLIBRI hue via a hand-crafted multilingual lexicon. Many borderline terms (salad/lime/swamp; mustard/beige/golden/brown; pink/coral/burgundy) are assigned by author judgment with no inter-annotator reliability, alternative mapping, or sensitivity analysis. Because yellow attracts more such borderline terms, systematic over-assignment away from yellow would mechanically produce the reported Green > Red > Yellow order. The central ranking cannot be distinguished from a mapping artifact without at least (i) dual independent annotation of a response sample and (ii) a leave-one-mapping-rule-out or coarser/finer lexicon reanalysis.
  2. Table III and §IV: the headline percentages (65.55%, 45.87%, 39.06%) are reported without confidence intervals, bootstrap uncertainty, or any inferential test of the Green > Red > Yellow ordering (or of saturation/intensity effects in Fig. 4). With ~320 responses per hue, such tests are feasible; without them the claim that categories “differ in their consistency” remains descriptive only and is not yet load-bearing evidence for unequal category breadth.
  3. §III Methodology: stimuli were shown in an online form with no display calibration, white-point control, or lighting instructions. For a color-perception claim about saturation/intensity robustness, uncontrolled sRGB rendering and mixed devices can systematically shift low-saturation and medium-intensity samples (especially yellows) toward brown/gray/beige—the very alternatives that lower yellow consistency. Either restrict claims to “naming under typical web viewing” or add a controlled lab/replication subset; as written, the perceptual-stability interpretation overreaches the uncontrolled setup.
  4. Table I and stimulus selection: all 18 samples come from the authors’ own COLIBRI fuzzy model (self-cited), and consistency is recovery of the COLIBRI hue label. Free responses supply independent data, so this is not pure circularity, but the grid may over-sample regions where COLIBRI already treats green as broad and yellow as narrow. A brief comparison to an independent atlas (e.g., Munsell or WCS-style chips) or an explicit statement that results are COLIBRI-conditioned would clarify external validity of the “broader perceptual region” claim.

Circularity Check

1 steps flagged · score 2.0 of 10

Mild self-citation of authors' COLIBRI model for stimulus selection; free-naming responses and consistency percentages remain independent empirical measurements, not forced by construction.

  1. self citation load bearing [Section III Methodology; Table I; Algorithm 1; citation [2]]
    "A set of 18 color samples was selected from the COLIBRI dataset to represent different shades of these colors. ... The stimuli (n=18) for the experiment were selected from the COLIBRI dataset [2]. ... Ch ← 1/|Dh| ∑ I(ci = hi)"

    Stimuli hues hi are taken from the authors' own COLIBRI model (self-cited [2]). Consistency is then the rate at which free responses map back to those same hi labels. This creates a mild dependence of the measured quantity on the authors' prior labeling scheme. However, free responses are independent and frequently map elsewhere, so the percentages are not forced by definition; the self-citation is not fully load-bearing for the ranking claim.

full rationale

The paper is an empirical free color-naming study, not a first-principles derivation. Stimuli (18 samples) are taken from the authors' prior COLIBRI fuzzy model (self-cited as [2]), with hues labeled red/yellow/green by that model (Table I). Consistency is then defined as the fraction of free multilingual responses that, after author-constructed lexicon mapping (Algorithm 1 steps 5–6, Table II), equal the original COLIBRI hue label. Free participant responses (n=92) are independent data that can and do map to many non-matching categories (salad/lime for green, pink/coral for red, brown/mustard/beige for yellow; Table III, word clouds, heatmaps). The reported ranking Green 65.55% > Red 45.87% > Yellow 39.06% is therefore not equivalent to the COLIBRI labels by construction; it is an observed frequency. The hand-built mapping dictionary is a methodological choice that could introduce bias (a validity concern), but it does not make the percentages tautological. No fitted parameters are re-labeled as predictions, no uniqueness theorem is imported, and no ansatz is smuggled. Self-citation of COLIBRI supplies the stimulus set but is not load-bearing for the central empirical claim. Score 2 reflects only that minor self-citation; the derivation chain is otherwise self-contained against the collected free-naming data.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

Empirical free-naming study. The main free parameters are the hand-crafted multilingual lexicon and the discrete stimulus grid chosen from COLIBRI. Core axioms are standard psychophysics assumptions plus the unstated claim that uncontrolled online displays do not systematically distort naming. No new physical entities are postulated.

free parameters (2)
  • dictionary category mappings = hand-crafted Table II
    Authors define which free strings (including food, plant, and shade terms) map to which category; these choices directly determine the reported consistency percentages.
  • stimulus selection grid = IDs 4,5,7,8,10,11,22,23,25,26,28,29,31,32,34,35,37,38
    Choice of the 18 specific COLIBRI IDs spanning low/medium/high saturation and medium/light intensity defines the appearance variation that is tested.
assumptions (3)
  • domain assumption Free color names, after lexicon mapping, reflect perceptual category membership
    Core premise of the free-naming paradigm; invoked throughout Results and Algorithm 1.
  • domain assumption COLIBRI hue labels correctly identify the intended red/yellow/green regions of color space
    Stimuli are selected and consistency is scored against these labels (Table I, Algorithm 1).
  • ad hoc to paper Online uncontrolled displays and mixed lighting do not systematically distort color naming
    Experiment run via online form at a conference; no display calibration or viewing-condition control is described (Section III).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Not All Color Categories Are Equally Stable: A Multilingual Free Color Naming Experiment." pith.science (2026). https://pith.science/paper/MK26LGF3

@misc{pith2026260710465,
  author       = {Pith},
  title        = {Pith review of: Not All Color Categories Are Equally Stable: A Multilingual Free Color Naming Experiment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MK26LGF3}},
  note         = {Machine review of arXiv:2607.10465}
}
read the original abstract

Color naming is an important part of human color perception. Its task is to allow people to describe continuous colors using discrete color categories. However, the boundaries between color categories are often unclear, and some colors may be perceived differently depending on their saturation and brightness. While certain color categories remain recognizable across a wide range of shades, others may be associated with different color names when their appearance changes. This study investigates the consistency of color naming for red, yellow, and green color categories using a free color-naming experiment. A set of 18 color samples was selected from the COLIBRI dataset to represent different shades of these colors. Participants (n = 92) were asked to freely assign color names to each sample in Kazakh, Russian, or English without being limited to predefined categories. The results show that color categories differ in their consistency. Green shades were consistently identified as green despite variations in appearance, whereas yellow shades received a wider variety of names, including gold- and brown-related descriptions. Red shades showed moderate naming consistency. Our findings suggest that some color categories occupy broader perceptual regions than others and may therefore be more robust to visual variations. The study results can be used to develop perceptually meaningful color models and color naming systems.

Figures

Figures reproduced from arXiv: 2607.10465 by the authors.

Figure 1
Figure 1. Experimental workflow for evaluating color category consistency using multilingual free color-naming responses [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Sample from the experiment suggest that color categories may exhibit different levels of perceptual stability. However, the extent to which entire color categories preserve naming consistency under variations in appearance remains insufficiently understood. Although many studies have examined color categorization, focal colors, and cross-linguistic color naming, limited at￾tention has been paid to comparing the nami… view at source ↗
Figure 3
Figure 3. Word clouds of color terms for RED, YELLOW, and GREEN colors [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Consistency scores by hue, saturation, and intensity [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Percentage heatmap of mapped color categories for [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

31 extracted references · 1 linked inside Pith

  1. [1]

    What we talk about when we talk about colors,

    C. R. T womey, G. Roberts, D. H. Brainard, and J. B. Plotkin, “What we talk about when we talk about colors,” Proceedings of the National Academy of Sciences , vol. 118, no. 39, 2021. [Online]. Available: http://dx.doi.org/10.1073/pnas.2109237118

  2. [2]

    Colibri fuzzy model: Color linguistic-based representation and interpretation,

    P . Shamoi, N. Toganas, M. Muratbekova, E. Kadyrgali, A. Y erkin, A. Igali, M. Ziyada, A. Adilova, A. Karatayev, and Y . Torekhan, “Colibri fuzzy model: Color linguistic-based representation and interpretation,” IEEE Access , vol. 13, pp. 205 932–205 956, 2025

  3. [3]

    Fuzzy color spaces: A conceptual approach to color vision,

    J. Chamorro-Martínez, J. M. Soto-Hidalgo, P . Martínez-Jiménez, and D. Sánchez, “Fuzzy color spaces: A conceptual approach to color vision,” IEEE Transactions on Fuzzy Systems , vol. 25, pp. 1264–1280, 2017

  4. [4]

    Variation of saturation across hue affects unique and typical hue choices,

    C. Witzel, “Variation of saturation across hue affects unique and typical hue choices,” i-Perception, vol. 10, 2019

  5. [5]

    Focal colors as perceptual anchors of color categories,

    C. Witzel, J. Maule, and A. Franklin, “Focal colors as perceptual anchors of color categories,” Journal of Vision , vol. 13, pp. 1164–1164, 2013

  6. [6]

    Berlin and P

    B. Berlin and P . Kay, Basic Color Terms: Their Universality and Evolution. University of California Press, 1969

  7. [7]

    Coherence of achromatic, primary and basic classes of colour categories,

    D. Mylonas and L. D. Griffin, “Coherence of achromatic, primary and basic classes of colour categories,” Vision Research, vol. 175, pp. 14–22, 2020

  8. [8]

    The color lexicon of american english

    D. Lindsey and A. M. Brown, “The color lexicon of american english.” Journal of vision , vol. 14 2, 2014

Show all 31 references
  1. [9]

    Universality of color names,

    D. T. Lindsey and A. M. Brown, “Universality of color names,” Proceedings of the National Academy of Sciences , vol. 103, pp. 16 608– 16 613, 2006

  2. [10]

    Color naming across languages,

    P . Kay and B. Berlin, “Color naming across languages,” Language, 1997

  3. [11]

    Russian blues reveal effects of language on color discrimination,

    J. e. a. Winawer, “Russian blues reveal effects of language on color discrimination,” PNAS, 2007

  4. [12]

    Toward a universal color naming system: A clustering-based approach using multisource data,

    A. Sabitkyzy, M. Shagyrov, and P . Shamoi, “Toward a universal color naming system: A clustering-based approach using multisource data,”

  5. [13]

    Available: https://arxiv.org/abs/2604.03235

    [Online]. Available: https://arxiv.org/abs/2604.03235

  6. [14]

    Color perception: Objects, con- stancy, and categories,

    C. Witzel and K. R. Gegenfurtner, “Color perception: Objects, con- stancy, and categories,” Annual Review of Vision Science , vol. 4, pp. 475–499, 2018

  7. [15]

    Color naming reflects optimal partitions of color space,

    T. Regier, P . Kay, and N. Khetarpal, “Color naming reflects optimal partitions of color space,” Proceedings of the National Academy of Sciences, vol. 104, no. 4, pp. 1436–1441, 2007

  8. [16]

    Fuzzy color space for apparel coordination,

    P . Shamoi, A. Inoue, and H. Kawanaka, “Fuzzy color space for apparel coordination,” Open Journal of Information Systems (OJIS) , vol. 1, pp. 20–28, 01 2014

  9. [17]

    Focal colors are universal after all,

    T. Regier, P . Kay, and R. S. Cook, “Focal colors are universal after all,” Proceedings of the National Academy of Sciences , vol. 102, no. 23, pp. 8386–8391, 2005

  10. [18]

    Focal colors across languages are representative members of color categories,

    J. T. Abbott, T. L. Griffiths, and T. Regier, “Focal colors across languages are representative members of color categories,” Proceedings of the National Academy of Sciences , vol. 113, no. 40, pp. 11 178– 11 183, 2016

  11. [19]

    Invariant categorical color regions across illuminant change coincide with focal colors,

    T. Morimoto, Y . Y amauchi, and K. Uchikawa, “Invariant categorical color regions across illuminant change coincide with focal colors,” Journal of Vision , vol. 22, no. 14, p. 3208, 2022

  12. [20]

    Variations in normal color vision. VII. relationships between color naming and hue scaling,

    K. J. Emery, V . J. Volbrecht, D. H. Peterzell, and M. A. Webster, “Variations in normal color vision. VII. relationships between color naming and hue scaling,” Vision Research, vol. 141, pp. 66–75, 2017

  13. [21]

    The Sapir-Whorf hypothesis and probabilistic inference: Evidence from the domain of color,

    E. Cibelli, Y . Xu, J. L. Austerweil, T. L. Griffiths, and T. Regier, “The Sapir-Whorf hypothesis and probabilistic inference: Evidence from the domain of color,” PLoS ONE , vol. 11, no. 7, p. e0158725, 2016

  14. [22]

    Perceptual constraints on colours induce the universality of linguistic colour categorisation,

    T. Gong, H. Gao, Z. Wang, and L. Shuai, “Perceptual constraints on colours induce the universality of linguistic colour categorisation,” Scientific Reports, vol. 9, p. 7719, 2019

  15. [23]

    An empirical study of basic color terms in chinese,

    S. Dai and A. Z. Zainal, “An empirical study of basic color terms in chinese,” Forum for Linguistic Studies , vol. 7, no. 5, pp. 861–872, 2025

  16. [24]

    Can language models encode perceptual structure without grounding? A case study in color,

    M. Abdou, A. Kulmizev, D. Hershcovich, S. Frank, E. Pavlick, and A. Søgaard, “Can language models encode perceptual structure without grounding? A case study in color,” in Proceedings of the 25th Conference on Computational Natural Language Learning (CoNLL) . Association for C...

  17. [25]

    Colour categories are reflected in sensory stages of colour perception when stimulus issues are resolved,

    L. Forder, X. He, and A. Franklin, “Colour categories are reflected in sensory stages of colour perception when stimulus issues are resolved,” PLoS ONE , vol. 12, no. 5, p. e0178097, 2017

  18. [26]

    Comparative analysis of color models for human perception and visual color difference,

    A. Burambekova and P . Shamoi, “Comparative analysis of color models for human perception and visual color difference,” in 2025 IEEE 5th In- ternational Conference on Smart Information Systems and Technologies (SIST), 2025, pp. 1–6

  19. [27]

    Color models in image processing: a review and experimental comparison,

    M. Muratbekova, N. Toganas, A. Igali, M. Shagyrov, E. Kadyrgali, A. Y erkin, and P . Shamoi, “Color models in image processing: a review and experimental comparison,” Discover Applied Sciences , vol. 8, no. 5, p. 494, Mar 2026

  20. [28]

    Lakoff and M

    G. Lakoff and M. Johnson, Metaphors We Live By . University of Chicago Press, 1980

  21. [29]

    Grounded cognition,

    L. W . Barsalou, “Grounded cognition,” Annual Review of Psychology , 2008

  22. [30]

    Online colour naming experiment using Munsell colour samples,

    D. Mylonas and L. MacDonald, “Online colour naming experiment using Munsell colour samples,” in Proceedings of the 5th European Conference on Colour in Graphics, Imaging, and Vision (CGIV) . Joensuu, Finland: IS&T, 2010, pp. 27–32

  23. [31]

    An online color naming experiment in Russian using Munsell color samples,

    G. V . Paramei, Y . A. Griber, and D. Mylonas, “An online color naming experiment in Russian using Munsell color samples,” Color Research & Application, vol. 43, pp. 358–374, 2018

Pith tools

Reviewed July 14, 2026 · model on record in the stance chip above.