REVIEW 4 major objections 4 minor 2 cited by
Understanding Bias in Perceiving Dimensionality Reduction Projections
T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Visual interestingness biases practitioners' selection of dimensionality-reduction projections over faithfulness, and color-encoded labels intensify the bias.
desk verdict Creative adversarial study design, but the synthetic faithfulness scores need a manipulation check before the 'over faithfulness' claim holds. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The experimental machinery has three load-bearing parts. First, an active ranking algorithm converts pairwise "which is more visually interesting?" choices into stable rankings of 20 projections per condition with minimal trials. Second, the adversarial faithfulness-score generator assigns random scores in [0,1] so that the less interesting projection wins on three or four of five metrics, creating a direct conflict between what looks good and what the numbers say. Third, Spearman rank correlations between the visual-interest rankings and the analytical-preference rankings, analyzed with two-way ANOVA, quantify the bias; dual-system theory supplies the explanatory frame, mapping visual appeal to System 1 and score-based evaluation to System 2.
What would settle it
A replication that replaces the synthetic scores with real, verified faithfulness metrics for the same projections and tells participants the scores are trustworthy; if preferences then track faithfulness rather than visual interest, the original result depended on distrust of the numbers, not on visual bias. A cheaper check is to add a post-task question asking participants how much they believed the synthetic scores predicted true quality.
Extended reading notes
Core claim
The central discovery is that visual interestingness—the degree to which a projection shows salient, distinctive, or aesthetically appealing patterns—dominates analytical preference for DR projections even when faithfulness scores explicitly favor a different projection. The authors created pairwise comparisons where five synthetic metrics, labeled A through E, were randomly generated so that three or four of them gave higher scores to the less visually interesting projection. Across 16 combinations of color encoding and exposure time, the Spearman correlation between visual-interest rankings and analytical-preference rankings stayed positive (around 0.25), and a two-way ANOVA found a significant effect of color encoding (F1,336 = 65.10, p < .001) but no significant effect of exposure time. Interviews confirmed that participants relied on visual appearance, felt that color-coded projections drew more attention, and were largely unaware of the bias. The paper interprets this through dual-system theory: visual interest triggers fast System 1 processing, while faithfulness requires slower System 2 reasoning that is often skipped.
Load-bearing premise
In Phase 2, the claim rests on participants accepting five randomly generated numbers, labeled only "metrics A through E" and rigged to favor less interesting projections, as credible indicators of true faithfulness; if they treated those numbers as meaningless or fake, the experiment would show a preference for attractive plots over arbitrary numbers, not over actual faithfulness.
Editorial extensions
If this is right
- When faithfulness scores conflict with visual appeal, practitioners tend to follow the visual appeal, so projection-selection interfaces that only show metric values may fail to steer analysts toward faithful embeddings.
- Color-encoding class labels amplifies the bias, so monochrome displays or shape encodings are plausible mitigations, as the paper recommends.
- Merely giving analysts more time may not remove the bias, since participants in the study decided in about five seconds even when allowed fifteen.
- Projections with well-separated classes and clumped clusters are perceived as visually interesting, so these visual features are the ones most likely to trigger the bias.
- The paper's proposed mitigation strategies—deactivating System 1 through visual simplification, making faithfulness scores visually salient, and improving DR literacy—are direct consequences of the identified mechanism and remain to be evaluated.
Reading between the lines
- The same perceptual mechanism likely extends beyond DR projection choice to any visual analytics setting where a salient, attractive layout competes with an objective quality score, so the bias may be a general visualization phenomenon rather than DR-specific.
- The adversarial-score design could be reused as a calibration instrument: presenting the same projections with real, verified faithfulness metrics would test whether the bias persists when the numbers are known to be trustworthy.
- Because participants were unaware of the bias and some denied it, a practical next step is to measure whether simply warning analysts about the bias, or requiring them to justify their choice, changes preferences.
- Individual-level analyses could test whether DR literacy moderates the effect; the paper reports bias across self-reported literacy levels but does not statistically compare subgroups.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a two-phase user study testing whether practitioners' analytical preference for dimensionality reduction (DR) projections is biased toward visual interestingness over faithfulness. In Phase 1, 16 participants produced visual-interest rankings of projections via active pairwise comparisons; in Phase 2, 16 different participants chose projections for cluster analysis while viewing five artificially generated, adversarially assigned 'faithfulness' scores. The authors test three hypotheses: H1 that visual interestingness more strongly influences analytical preference than faithfulness, H2 that color encoding intensifies the bias, and H3 that shorter exposure time intensifies the bias. The quantitative result is a significant effect of color encoding on the correlation between visual-interest and analytical-preference rankings (F1,336 = 65.10, p < .001) and no significant effect of exposure time (F1,336 = 0.22, p = 0.64). The paper also reports a regression analysis of visual features and qualitative interview findings, and proposes mitigation strategies.
Significance. If the central claim is upheld after revision, the paper addresses an important and timely problem for visualization and visual analytics: users may choose DR projections for perceptual appeal rather than structural fidelity. The adversarial score assignment is a genuinely useful design because it makes a null result possible, and the positive correlation plus the color-encoding effect are promising evidence. The qualitative findings on bias awareness and the proposed mitigation strategies are also valuable. However, the central inference depends on participants accepting fabricated faithfulness scores as credible, which is not checked, and the paper currently overstates H3 in the abstract. These issues are fixable within the manuscript's scope and do not require rejecting the work.
major comments (4)
- [Abstract and §4.1] The abstract states that the bias 'intensifies with color-encoded labels and shorter exposure time,' but the quantitative analysis in §4.1 reports no significant effect of EXPOSURE TIME (F1,336 = 0.22, p = 0.64) and no interaction. The qualitative Finding 2 in §5, based on ten self-reports, is explicitly called 'weak support' by the authors. The shorter-exposure claim must be removed from the abstract or explicitly qualified as qualitative-only, and the discussion should reconcile the quantitative null result with the qualitative reports.
- [§3.2.2 (Generating and presenting faithfulness scores)] The Phase 2 manipulation relies on participants treating five arbitrary numbers in [0,1], labeled only 'metrics A through E' and rigged to favor less visually interesting projections, as credible indicators of projection faithfulness. The paper provides no manipulation check, no comprehension check, and no pilot evidence that participants understood these scores as genuine faithfulness information. Without such a check, the positive correlation between visual-interest rankings and analytical-preference rankings could reflect a preference for visually interesting plots over arbitrary numeric displays rather than a bias against faithfulness. This is load-bearing for H1, so the authors should either add a manipulation check or a post-task credibility rating, or substantially reframe the claim to avoid equating the synthetic scores with faithfulness.
- [§4.1 (Analysis design)] The statistical reporting is internally inconsistent. The text says there are 16 combinations (2 color × 2 exposure × 4 dataset variations), each contributing 4 × 4 = 16 Spearman correlations, which yields 256 observations, but the reported ANOVA degrees of freedom (F1,336) imply 340 observations after four model parameters. Please clarify the exact unit of analysis, report how Phase 1 and Phase 2 rankings were paired despite having different participants, and re-run the analysis at the correct unit or explain the discrepancy.
- [§4.1 (Results and discussions)] The paper reports only that correlations 'range around 0.25' and gives F and p values for the ANOVA, but it never reports the mean, standard deviation, or confidence intervals for the Spearman correlations in each condition, nor the effect size for the color effect. Since the central conclusion is that the correlation is positive and larger with color encoding, the authors should report these descriptive statistics and effect sizes explicitly, including the values shown only in Figure 3.
minor comments (4)
- [§4.1] The sentence 'We find a significant effect on COLOR ENCODING ... confirming H1' appears to conflate H1 with H2; the color effect supports H2, while H1 is supported by the positive correlation itself. Please reword to avoid confusing the reader.
- [§3.2.1 (Stimuli)] The stimuli-sampling procedure is deferred to 'Appendix A,' but no appendix is present in the submitted manuscript. Please include the appendix or add a brief in-text description so the stratified sampling of datasets and projections is reproducible.
- [§5 (Finding 2)] The qualitative evidence for the exposure-time effect is based on self-reports in an interview setting where participants may be primed by the study design; the text already calls this 'weak support,' but the authors should avoid using it to endorse H3 in the conclusion and abstract.
- [§3.2.2] Please clarify whether the 'five pairs of faithfulness scores' means five scores per projection or five paired comparisons per trial; the current wording is ambiguous and the reader cannot determine how the scores were visually arranged for the participant.
Circularity Check
No significant circularity: the bias claim rests on separately measured rankings and an adversarial score assignment that permits a null result.
full rationale
The paper is an empirical user study rather than a derivation. Visual-interest rankings (Phase 1) and analytical-preference rankings (Phase 2) are elicited from different participants with different questions, and the faithfulness scores are artificially generated to favor the less interesting projection ('we assign higher scores to the less visually interesting projection on three or four randomly selected metrics'). This adversarial construction makes a negative or zero correlation possible, so the observed positive Spearman correlation is not forced by construction. The central inference therefore does not reduce to its inputs. Authors' prior work is cited for auxiliary materials (the 96-dataset pool [13], the ZADU metric library [14], CLAMS cluster-count estimation [17]) and for discussion of t-SNE/UMAP misuse [16,15]; these are external published artifacts used as tools or background, not as the evidence establishing the bias, so they do not constitute load-bearing self-citation. The only substantive threat—that participants may not have treated the fake A–E scores as credible faithfulness information—is a manipulation-check/construct-validity concern, not a circularity, because the outcome measure was not defined in terms of, nor fitted to, the predictor. No equation or fitted parameter is renamed as a prediction.
Assumptions & free parameters
free parameters (3)
- Faithfulness score generation rule =
3 or 4 of 5 metrics favor the less interesting projection; values uniform in [0,1]
- Visual interestingness score transformation =
score = 21 - rank
- Exposure time levels =
7 s and 15 s
assumptions (4)
- domain assumption Dual-system theory (System 1 and System 2) describes the cognitive processes at play in DR projection selection.
- domain assumption Artificial metrics A-E are accepted by participants as faithful indicators of projection quality.
- domain assumption The active ranking algorithm produces stable and valid rankings from 50 pairwise comparisons per 20 projections.
- domain assumption University students with scatterplot experience are representative of practitioners.
Cite this review
Pith. "Pith review of Understanding Bias in Perceiving Dimensionality Reduction Projections." pith.science (2026). https://pith.science/paper/YKHF2OVN
@misc{pith2026250720805,
author = {Pith},
title = {Pith review of: Understanding Bias in Perceiving Dimensionality Reduction Projections},
year = {2026},
howpublished = {\url{https://pith.science/paper/YKHF2OVN}},
note = {Machine review of arXiv:2507.20805}
}
read the original abstract
Selecting the dimensionality reduction technique that faithfully represents the structure is essential for reliable visual communication and analytics. In reality, however, practitioners favor projections for other attractions, such as aesthetics and visual saliency, over the projection's structural faithfulness, a bias we define as visual interestingness. In this research, we conduct a user study that (1) verifies the existence of such bias and (2) explains why the bias exists. Our study suggests that visual interestingness biases practitioners' preferences when selecting projections for analysis, and this bias intensifies with color-encoded labels and shorter exposure time. Based on our findings, we discuss strategies to mitigate bias in perceiving and interpreting DR projections.
Figures
Forward citations
Cited by 2 Pith papers
-
Stop Misusing t-SNE and UMAP for Visual Analytics
A review of 136 visual analytics papers and two interview studies show t-SNE and UMAP are frequently used for tasks they cannot support, and that practitioner literacy issues explain the persistent misuse.
-
FlexMUSE: Multimodal Unification and Semantics Enhancement Framework with Flexible interaction for Creative Writing
FlexMUSE, a claimed multimodal creative-writing framework and its ArtMUSE dataset, are unsupported because the submitted full text is an unrelated dimensionality-reduction paper (UMATO).
Reference graph
Works this paper leans on
-
[1]
M. M. Abbas, M. Aupetit, M. Sedlmair, and H. Bensmail. Clustme: A visual quality measure for ranking monochrome scatterplots based on cluster patterns. Computer Graphics Forum, 38(3):225–236, 2019. doi: 10.1111/cgf.13684 3
-
[2]
S. S. Bae, T. Fujiwara, C. Tseng, and D. Szafir. Uncovering how scat- terplot features skew visual class separation. In ACM CHI, 2025. 4
work page 2025
-
[3]
J. Bernard, M. Hutter, M. Zeppelzauer, D. Fellner, and M. Sedlmair. Comparing visual-interactive labeling with active learning: An exper- imental study. IEEE Transactions on Visualization and Computer Graphics, 24(1):298–308, 2018. doi: 10.1109/TVCG.2017.2744818 5
-
[4]
A. Bibal and B. Fr ´enay. Learning interpretability for visualizations using adapted cox models through a user experiment, 2016. 2, 4
work page 2016
-
[5]
D. Cashman, M. Keller, H. Jeon, B. C. Kwon, and Q. Wang. A criti- cal analysis of the usage of dimensionality reduction in four domains. IEEE Transactions on Visualization and Computer Graphics , pp. 1– 20, 2025. doi: 10.1109/TVCG.2025.3567989 1, 2
-
[6]
A. Colin Cameron and F. A. Windmeijer. An r-squared measure of goodness of fit for some common nonlinear regression models. Jour- nal of Econometrics, 77(2):329–342, 1997. doi:10.1016/S0304-4076(96) 01818-0 4
-
[7]
M. Correll and M. Gleicher. Error bars considered harmful: Exploring alternate encodings for mean and error. IEEE Transactions on Visu- alization and Computer Graphics , 20(12):2142–2151, 2014. doi: 10. 1109/TVCG.2014.2346298 5
arXiv 2014
-
[8]
J. H. Friedman. Exploratory projection pursuit. Journal of the Ameri- can statistical association, 82(397):249–266, 1987. 1, 2
work page 1987
Show all 39 references
-
[9]
Goffin, J
P. Goffin, J. Boy, W. Willett, and P. Isenberg. An exploratory study of word-scale graphics in data-rich text documents. IEEE Transactions on Visualization and Computer Graphics , 23(10):2275–2287, 2017. doi: 10.1109/TVCG.2016.2618797 5
2017
-
[10]
T. M. Green, W. Ribarsky, and B. Fisher. Visual analytics for complex concepts using a human cognition model. In 2008 IEEE Symposium on Visual Analytics Science and Technology, pp. 91–98, 2008. doi: 10. 1109/VAST.2008.4677361 2, 5
2008
-
[11]
Healey and J
C. Healey and J. Enns. Attention and visual memory in visualization and computer graphics. IEEE Transactions on Visualization and Com- puter Graphics, 18(7):1170–1188, 2012. doi: 10.1109/TVCG.2011.127 2, 3
2012 doi
-
[12]
K. G. Jamieson and R. Nowak. Active ranking using pairwise compar- isons. Advances in neural information processing systems , 24, 2011. 3, 4
2011
-
[13]
H. Jeon, M. Aupetit, D. Shin, A. Cho, S. Park, and J. Seo. Measuring the validity of clustering validation datasets. IEEE Transactions on Pattern Analysis and Machine Intelligence , 47(6):5045–5058, 2025. doi: 10.1109/TPAMI.2025.3548011 3
2025
-
[14]
H. Jeon, A. Cho, J. Jang, S. Lee, J. Hyun, H.-K. Ko, J. Jo, and J. Seo. Zadu: A python library for evaluating the reliability of dimensionality reduction embeddings. In 2023 IEEE Visualization and Visual Ana- lytics (VIS), pp. 196–200. doi: 10.1109/VIS54172.2023.00048 2, 4
2023
-
[15]
H. Jeon, H. Lee, Y .-H. Kuo, T. Yang, D. Archambault, S. Ko, T. Fuji- wara, K.-L. Ma, and J. Seo. Unveiling high-dimensional backstage: A survey for reliable visual analytics with dimensionality reduction. In Proceedings of the 2025 CHI Conference on Human Factors in Com- puti...
2025
-
[16]
H. Jeon, J. Park, S. Shin, and J. Seo. Stop misusing t-sne and umap for visual analytics, 2025. 1, 2, 5
2025
-
[17]
H. Jeon, G. J. Quadri, H. Lee, P. Rosen, D. A. Szafir, and J. Seo. Clams: A cluster ambiguity measure for estimating perceptual vari- ability in visual clustering. IEEE Transactions on Visualization and Computer Graphics, 30(1):770–780, 2024. doi: 10.1109/TVCG.2023 .3327201 2, 3, 4
2024 doi
-
[18]
P. Joia, D. Coimbra, J. A. Cuminato, F. V . Paulovich, and L. G. Nonato. Local affine multidimensional projection. IEEE Transactions on Vi- sualization and Computer Graphics, 17(12):2563–2571, 2011. doi: 10 .1109/TVCG.2011.220 4
2011
-
[19]
Kahneman
D. Kahneman. Thinking, fast and slow. 2011. 2
2011
-
[20]
Kobak and G
D. Kobak and G. C. Linderman. Initialization is critical for preserving global data structure in both t-sne and umap. Nature biotechnology, 39(2):156–157, 2021. doi: 10.1038/s41587-020-00809-z5
2021 doi
-
[21]
B. C. Kwon, B. Eysenbach, J. Verma, K. Ng, C. De Filippi, W. F. Stewart, and A. Perer. Clustervision: Visual supervision of unsuper- vised clustering. IEEE Transactions on Visualization and Computer Graphics, 24(1):142–151, 2018. doi: 10.1109/TVCG.2017.2745085 2, 3, 4
2018
-
[22]
L. v. d. Maaten and G. Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(Nov):2579–2605, 2008. 5
2008
-
[23]
McInnes, J
L. McInnes, J. Healy, and J. Melville. Umap: Uniform manifold ap- proximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426, 2018. 5
2018 arXiv
-
[24]
Morariu, A
C. Morariu, A. Bibal, R. Cutura, B. Fr ´enay, and M. Sedlmair. Pre- dicting user preferences of dimensionality reduction embedding qual- ity. IEEE Transactions on Visualization and Computer Graphics , 29(1):745–755, 2023. doi: 10.1109/TVCG.2022.3209449 2, 4
2023
-
[25]
Nguyen, P
Q. Nguyen, P. Eades, and S.-H. Hong. On the faithfulness of graph visualizations. In 2013 IEEE Pacific Visualization Symposium (Paci- ficVis), pp. 209–216, 2013. doi: 10.1109/PacificVis.2013.6596147 1, 2
2013
-
[26]
L. G. Nonato and M. Aupetit. Multidimensional projection for visual analytics: Linking techniques with distortions, tasks, and layout en- richment. IEEE Transactions on Visualization and Computer Graph- ics, 25(8):2650–2673, 2019. doi: 10.1109/TVCG.2018.2846735 1, 2
2019
-
[27]
A. V . Pandey, J. Krause, C. Felix, J. Boy, and E. Bertini. Towards understanding human similarity perception in the analysis of large sets of scatter plots. InProceedings of the 2016 CHI Conference on Human Factors in Computing Systems, CHI ’16, p. 3659–3669, 2016. doi: 10. 1...
2016
-
[28]
Seo and B
J. Seo and B. Shneiderman. A rank-by-feature framework for interac- tive exploration of multidimensional data. Information visualization, 4(2):96–113, 2005. 1, 2
2005
-
[29]
South, D
L. South, D. Saffo, O. Vitek, C. Dunne, and M. A. Borkin. Effective use of likert scales in visualization evaluations: A systematic review. Computer Graphics Forum, 41(3):43–55, 2022. doi: 10.1111/cgf.14521 3
2022 doi
-
[30]
Strobelt, D
H. Strobelt, D. Oelke, B. C. Kwon, T. Schreck, and H. Pfister. Guide- lines for effective usage of text highlighting techniques. IEEE Trans- actions on Visualization and Computer Graphics , 22(1):489–498,
-
[31]
Tseng, A
C. Tseng, A. Z. Wang, G. J. Quadri, and D. A. Szafir. Shape it up: An empirically grounded approach for designing shape palettes. IEEE Transactions on Visualization and Computer Graphics , 31(1):349– 359, 2025. doi: 10.1109/TVCG.2024.3456385 5
2025
-
[32]
Tversky and D
A. Tversky and D. Kahneman. Judgment under uncertainty: Heuristics and biases: Biases in judgments reveal some heuristics of thinking under uncertainty. science, 185(4157):1124–1131, 1974. 2
1974
-
[33]
A. C. Valdez, M. Ziefle, and M. Sedlmair. Priming and anchoring ef- fects in visualization. IEEE Transactions on Visualization and Com- puter Graphics, 24(1):584–594, 2018. doi: 10.1109/TVCG.2017.2744138 3
2018
-
[34]
Venna and S
J. Venna and S. Kaski. Local multidimensional scaling. Neural Net- works, 19(6):889–899, 2006. Advances in Self Organising Maps - WSOM’05. doi: 10.1016/j.neunet.2006.05.014 2
2006 doi
-
[35]
Y . Wang, K. Feng, X. Chu, J. Zhang, C.-W. Fu, M. Sedlmair, X. Yu, and B. Chen. A perception-driven approach to supervised dimension- ality reduction for visualization. IEEE Transactions on Visualization and Computer Graphics , 24(5):1828–1840, 2018. doi: 10.1109/TVCG. 2017.27...
2018
-
[36]
Wattenberg, F
M. Wattenberg, F. Vi´egas, and I. Johnson. How to use t-sne effectively. Distill, 2016. doi: 10.23915/distill.00002 2, 5
2016 doi
-
[37]
Wilkinson, A
L. Wilkinson, A. Anand, and R. Grossman. Graph-theoretic scagnos- tics. In Information visualization, IEEE symposium on , pp. 21–21. IEEE Computer Society, 2005. 4
2005
-
[38]
J. Xia, Y . Zhang, J. Song, Y . Chen, Y . Wang, and S. Liu. Revisit- ing dimensionality reduction techniques for visual cluster analysis: An empirical study. IEEE Transactions on Visualization and Com- puter Graphics, 28(1):529–539, 2022. doi: 10.1109/TVCG.2021.3114694 2, 3
2022
-
[2016]
doi: 10.1109/TVCG.2015.2467759 5
2015
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.