Pith. sign in

REVIEW 1 major objections 6 minor 70 references

Sketches outperform words at revealing cultural differences in thought

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · glm-5.2

2026-07-09 15:34 UTC pith:H6I2WMPQ

load-bearing objection Large-scale sketch analysis shows image-based concept representations capture cultural variation that translated word embeddings miss, but the language baseline may be too weak to support the compression claim the 1 major comments →

arxiv 2607.07267 v1 pith:H6I2WMPQ submitted 2026-07-08 cs.CY cs.CLphysics.soc-ph

Billions of Sketches Reveal Hidden Cultural Variation in Human Concepts

classification cs.CY cs.CLphysics.soc-ph
keywords conceptsculturalacrossconceptualsketchesvariationvisualhuman
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks whether human concepts are as universal as language-based analyses suggest. Words compress experience into shared labels, potentially hiding real differences in how people mentally represent everyday things. Using 2.6 billion sketches from 236 countries, the authors show that single concepts like pizza or phone split into multiple distinct visual forms rather than converging on one prototype. These visual representations diverge sharply from word-based semantic structure: concepts that look alike are rarely semantically close, and vice versa. Most consequentially, cross-cultural patterns in how people draw concepts align 45% more closely with independently measured cultural distances than do cross-cultural patterns in word meanings. The paper's central claim is that the modality through which we measure concepts determines how universal or diverse they appear, and that visual sketching exposes cultural and embodied variation that language compresses away.

Core claim

When people across 236 countries sketch everyday concepts, their drawings organize into stable visual exemplar clusters whose cross-cultural similarity patterns track established cultural distance measures 45% more closely than word-based similarity does. The geometry of these visual embedding spaces diverges systematically from word embedding spaces, showing that language and imagery encode different information about the same concepts. Concepts involving hands-on physical manipulation produce the most coherent visual clusters, linking conceptual structure to embodied experience rather than to dictionary definitions.

What carries the argument

The analysis pipeline embeds each sketch in a 384-dimensional latent space using a self-supervised vision transformer, reduces dimensions via PCA and UMAP, then clusters within each concept using DBSCAN with density-based validity optimization. Country-level visual similarity is computed via odds ratios measuring how strongly each country concentrates in specific visual clusters. This is compared against a word-embedding similarity network built from multilingual BERT embeddings of translated concept names, and against a cultural distance network derived from World Values Survey data. Network similarity is measured through edge Jaccard overlap, neighborhood Jaccard overlap, and NormalizedMut

Load-bearing premise

The entire argument depends on the image embedding model capturing genuine conceptual variation in what people draw, rather than low-level artifacts like stroke thickness, drawing speed, device type, or digital skill, all of which could correlate with country and produce spurious cultural clustering.

What would settle it

If the 45% cultural alignment advantage disappeared after controlling for drawing device type, interface differences, or socioeconomic factors that correlate with country, the central claim that visual representations preserve culturally meaningful structure would be substantially weakened.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If visual representations preserve cultural information that language compresses, then AI systems trained only on text will systematically miss culturally variable conceptual structure that matters for cross-cultural communication and design.
  • The finding that haptic concepts cluster most coherently suggests that embodied interaction shapes visual conceptualization, providing a testable bridge between sensorimotor experience and abstract concept structure.
  • Debates about conceptual universality that rely solely on linguistic data may be asking a modality-specific question and arriving at modality-specific answers, limiting the generality of both universalist and relativist claims.
  • Large-scale behavioral datasets from gamified platforms can serve as high-resolution probes of cognitive and cultural structure at population scale, complementing controlled laboratory studies.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the 45% alignment advantage of sketches over words holds after controlling for drawing device, internet access patterns, and English-language bias in the dataset, it would suggest that visual imagination is a more culturally sensitive channel than language for measuring conceptual diversity.
  • The divergence between visual and linguistic geometry raises the possibility that multilingual large language models, which learn from text alone, encode a systematically impoverished model of human conceptual structure, missing the exemplar-level variation that sketches preserve.
  • The clustering of haptic concepts into more coherent visual forms could imply that concepts grounded in shared physical manipulation are less culturally variable than concepts grounded in visual or social experience, though the paper claims the opposite direction, that haptic concepts cluster more, not less.
  • If sketch-based cultural networks align better with survey-based cultural distances than word-based networks do, then sketching behavior might serve as an unobtrusive, scalable proxy for cultural distance measurement that does not require self-report instruments.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 6 minor

Summary. This manuscript analyzes 2.6 billion sketches from the QuickDraw dataset across 236 countries to examine cultural variation in human conceptual structure. The authors find that concepts unfold into multiple visual exemplar clusters, that clusterability correlates with haptic sensorimotor properties, that image-based and word-based embedding geometries diverge, and that cross-cultural similarities derived from sketches align 45% more closely with World Values Survey cultural distances than do text-based measures. The central claim is that visual representations preserve rich semantic and cultural structure that language compresses, and that the modality of measurement critically affects conclusions about conceptual universality. The dataset is unprecedented in scale and the network comparison methodology is generally sound, with the cultural benchmark (WVS CFs index) being external to both sketch and word data. However, the headline 45% figure rests on a language baseline that may be structurally disadvantaged, and no statistical uncertainty is reported for this load-bearing comparison.

Significance. The paper's primary contribution is methodological and conceptual: it demonstrates that large-scale sketch data can serve as a high-resolution probe of cultural variation in conceptual structure, complementing word-based analyses. The scale of the dataset (2.6 billion sketches, 236 countries) is genuinely unprecedented for image-based cognitive research, and the pipeline (DINOv2 → PCA → UMAP → DBSCAN with DBCV optimization) is technically reasonable. The finding that haptic interaction correlates with clusterability is a novel, falsifiable empirical result linking embodied cognition to visual representation. The network comparison framework—comparing image- and language-based country similarity networks against an independent cultural benchmark—is well-constructed. Code availability and a 50M sketch sample for reproducibility are commendable. The paper also makes a timely contribution to debates about whether LLMs can capture human conceptual structure by showing that word-level representations mask latent visual variation.

major comments (1)
  1. The headline 45% cultural alignment figure (Abstract; Results, Fig. 4) is reported as an average across three network similarity metrics (edge similarity, neighborhood similarity, community overlap) and multiple edge-filtering thresholds (1%–90%), with no single confidence interval, standard error, or statistical test reported for the aggregate figure. The embedded histogram in Fig. 4 shows the distribution of ratios, but the reader cannot assess whether the 45% average is statistically distinguishable from a null hypothesis of no difference between image and language alignment. Given that this is the most prominent quantitative claim in the paper, a formal test (e.g., bootstrap CI on the ratio, or a permutation test comparing image-culture vs. language-culture similarity across thresholds) is needed. The robustness checks in Extended Data Fig. 13 (25-node and 50-node networks) are noted
minor comments (6)
  1. Figure 3: The word-based network is described as showing 'weaker and less coherent clustering,' but the figure caption does not specify the edge retention threshold or filtering method used for this visualization. Adding this information would make the visual comparison in Fig. 3 more rigorous.
  2. Methods, 'Word vs. Image Semantics': The macro average rank correlation of 0.098 is described as 'consistently low' across 'multiple correlation metrics and embedding models,' but the SI details referenced are not visible in the main text. A brief summary table or inline statistics would help readers evaluate this claim without consulting SI.
  3. The 21-cluster outlier for 'crow' is mentioned without explanation. A brief note on why this concept produces so many clusters would help contextualize the clustering distribution.
  4. Extended Data Table 1: The country distribution shows 41.3% US sketches. While the paper acknowledges this in the Discussion, the main text could note the degree of US overrepresentation earlier (e.g., in the dataset description in the Results) to set reader expectations.
  5. The paper uses LLaMA 3.3-70B to supplement concreteness and sensorimotor ratings for 38–40 concepts lacking crowd-sourced annotations. The agreement between LLM-generated and human ratings for the overlapping concepts is not reported. A brief validation (e.g., correlation between LLM and human scores for concepts with both) would strengthen confidence in the haptic-interaction finding.
  6. The phrase 'language models compress' (Abstract, Discussion) is used loosely. The paper compares word embeddings (Word2Vec, multilingual BERT) with image embeddings, not language models in the generative sense. Clarifying that 'language models' refers to embedding models would improve precision.

Circularity Check

0 steps flagged

No circularity found; derivation chain is self-contained against external benchmarks

full rationale

The paper's three main claims are each derived from independently constructed inputs compared against external benchmarks. (1) The clustering result (concepts unfold into multiple visual exemplars) is an empirical output of DBSCAN on DINOv2-embedded sketches — no step defines clusters in terms of the claimed outcome. (2) The visual-linguistic divergence (0.098 rank correlation) compares two independently generated embedding spaces (DINOv2 image embeddings vs. Word2Vec/multilingual BERT word embeddings) — neither is constructed from the other. (3) The 45% cultural alignment figure compares three independently constructed networks: an image-based country similarity network (odds ratios of country-cluster associations), a language-based network (cosine similarity of translated word embeddings), and a cultural benchmark (WVS CFs index from Muthukrishna et al. 2020, an external dataset). The comparison uses a configuration-model null preserving degree sequence. No network is defined in terms of another. The self-citation at ref 14 (Guilbeault, Baronchelli, Centola 2021) provides conceptual framing about word compression but is not load-bearing for any mathematical derivation. The skeptic's concern that the language baseline may be near-ceiling for concrete nouns is a construct-validity issue, not circularity — the language network is not defined in terms of the cultural distance, nor is the image network defined in terms of either. The derivation is self-contained.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 0 invented entities

No new entities, particles, forces, or dimensions are postulated. The paper introduces no new theoretical constructs beyond empirical observations. The free parameters are all data-driven pipeline choices rather than theoretical constants. The axioms are domain assumptions about the validity of the data and tools, not new postulates.

free parameters (6)
  • DBSCAN epsilon = optimized via DBCV
    Clustering radius parameter optimized per concept using density-based cluster validity; not a global constant but a data-driven parameter.
  • Noise threshold for clusterability = local minimum of bimodal noise distribution
    Post-hoc threshold determined from the data distribution to separate clusterable from non-clusterable concepts.
  • Grid density percentile (non-clusterable concepts) = 60th percentile
    Threshold for high-density grid cells, optimized against DBSCAN labels on clusterable concepts.
  • Network edge retention threshold = top 10%
    Retains top 10% strongest edges for community detection; robustness checked across 1-90%.
  • Odds ratio similarity band = 0.9-1.1
    Range within which country-cluster odds ratios are treated as zero similarity; chosen ad hoc.
  • PCA dimensions = 40
    Reduction from 384 DINOv2 dimensions to 40; no stated justification for this specific value.
axioms (5)
  • domain assumption DINOv2 image embeddings capture semantically meaningful structural variation in human sketches
    The entire analysis pipeline depends on the embedding space reflecting conceptual content rather than low-level visual artifacts. This is assumed, not independently validated for sketch data.
  • domain assumption QuickDraw sketches reflect genuine conceptual representation rather than task-specific drawing strategies
    The 20-second time limit and game format may elicit stereotyped or simplified drawings that do not fully represent conceptual structure. The paper assumes sketches are valid proxies for mental representation.
  • domain assumption Country-level IP address is a valid proxy for cultural context
    Users' countries are inferred from IP addresses. The paper assumes this captures meaningful cultural grouping, though it acknowledges demographic information is unavailable.
  • domain assumption Translated concept names in multilingual word embeddings provide a fair baseline for linguistic conceptual structure
    The word-vs-image comparison depends on translated concept names being representative of cross-linguistic conceptual structure. Translation artifacts could inflate the divergence between modalities.
  • standard math World Values Survey cultural distance is an independent ground truth for cultural similarity
    The CFs index from Muthukrishna et al. 2020 is treated as the cultural benchmark. This is a standard external dataset, but it itself has WEIRD biases and coverage gaps.

pith-pipeline@v1.1.0-glm · 21495 in / 3104 out tokens · 414249 ms · 2026-07-09T15:34:24.658098+00:00 · methodology

0 comments
read the original abstract

Claims about the universality of human concepts have been predominantly assessed through linguistic similarity across languages and cultures. However, words are effective as communication devices because they compress rich experiential variation into shared conventions, potentially obscuring hidden individual and cultural differences in how concepts are mentally represented. Here, we analyse 2.6 billion human-made sketches of common concepts from 236 countries and territories to examine conceptual structure through people's visual imagination. Consistent with recent work on image-based cognition, we find that single concepts unfold into multiple distinct visual exemplars, revealing latent information about similarities and differences in conceptual structure across cultures. This variation is strongest for concepts involving haptic interaction, suggesting that visual imagery reflects variation in embodied experience as much as conventional definitions. Comparing embedding models of sketches with word embedding models across languages, we find that their geometries diverge, with visual representations preserving rich semantic and cultural structure that language models compress. Cross-cultural similarities derived from sketches align 45% more closely with established cultural distances than do text-based measures. Together, these results suggest that patterns of human conceptual universality may depend critically on the modality through which concepts are measured, with large-scale sketching providing a direct, high-resolution probe of conceptual diversity across embodied and cultural dimensions of thought.

Figures

Figures reproduced from arXiv: 2607.07267 by Andrea Baronchelli, Arianna Pera, Douglas Guilbeault, Luca Maria Aiello, Mauro Martino, Nima Dehmamy.

Figure 1
Figure 1. Figure 1: Sketches of concepts provided by people around the world organize into distinct visual clusters. Visual clusters for six representative concepts (a–f) selected from the 344 available. Each point represents a drawing projected into a two-dimensional latent visual space. Points are colored according to their algorithmically assigned cluster, with gray points denoting drawings classified as random noise. Each… view at source ↗
Figure 2
Figure 2. Figure 2: Clusterability of sketches is selectively associated with haptic and sen￾sorimotor conceptual properties. Correlation between clusterability and conceptual properties of objects. Blue bars correspond to significant values (α = 0.05, after Bonfer￾roni correction), grey to non-significant ones. The density distribution of property scores is reported on the x -axis. To ensure that the correlations were not dr… view at source ↗
Figure 3
Figure 3. Figure 3: Image- and word-based concept networks exhibit divergent large-scale patterns of inter-cultural distances and clustering. Networks of countries based on sketch similarity (left) and word similarity (right). The top 100 nodes by number of sketches are shown. Colors denote structural communities of countries that share a high level of similarity with one another within each network, as identified by the Louv… view at source ↗
Figure 4
Figure 4. Figure 4: Cultural similarity aligns more closely with image-based than word￾based concept networks. Comparison between image- and language-based network sim￾ilarity with the cultural network across different metrics, expressed as ratios relative to a baseline defined by a null model. The networks are compared in terms of edge similarity, node similarity, and community overlap. Measurements are repeated on multiple … view at source ↗
Figure 5
Figure 5. Figure 5: Distribution of number of seconds passed before the sketched was recognized [PITH_FULL_IMAGE:figures/full_fig_p032_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Distributions of the percentage of noise in clustering (left) and the number of [PITH_FULL_IMAGE:figures/full_fig_p033_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Superposition of randomly sampled drawings within each cluster, for six represen [PITH_FULL_IMAGE:figures/full_fig_p034_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Steps of the grid-based pipeline on a selected example concepts ( [PITH_FULL_IMAGE:figures/full_fig_p034_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Performance of the alternative clustering methodology on clusterable concept cat [PITH_FULL_IMAGE:figures/full_fig_p035_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Similarity of rankings considering word and image embeddings, together with a [PITH_FULL_IMAGE:figures/full_fig_p036_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: NMI score for robustness of community structure across percentages of strongest [PITH_FULL_IMAGE:figures/full_fig_p037_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Comparison of network countries’ similarity based on image and words. Each [PITH_FULL_IMAGE:figures/full_fig_p040_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Comparison between image- and language-based network similarity with the cul [PITH_FULL_IMAGE:figures/full_fig_p041_13.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

70 extracted references · 70 canonical work pages · 4 internal anchors

  1. [1]

    & Kay, P

    Berlin, B. & Kay, P. Basic color terms: Their universality and evolution (Univ of Cali- fornia Press, 1969)

  2. [2]

    & Regier, T

    Kemp, C., Xu, Y. & Regier, T. Semantic typology and efficient communication. Annual Review of Linguistics 4, 109–128 (2018)

  3. [3]

    & Evans, J

    Lewis, M., Cahill, A., Madnani, N. & Evans, J. Local similarity and global variabil- ity characterize the semantic space of human languages. Proceedings of the National Academy of Sciences 120, e2300986120 (2023)

  4. [4]

    C., Jiang, S

    Xu, Y., Duong, K., Malt, B. C., Jiang, S. & Srinivasan, M. Conceptual relations predict colexification across languages. Cognition 201, 104280 (2020)

  5. [5]

    & Ran, Q

    Liang, Y., Xu, K. & Ran, Q. Shared structure of fundamental human experience revealed by polysemy network of basic vocabularies across languages. Scientific Reports 14, 5877 (2024)

  6. [6]

    & Tishby, N

    Zaslavsky, N., Kemp, C., Regier, T. & Tishby, N. Efficient compression in color naming and its evolution. Proceedings of the National Academy of Sciences 115, 7937–7942 (2018)

  7. [7]

    H., Norcliffe, E

    San Roque, L., Kendrick, K. H., Norcliffe, E. & Majid, A. Universal meaning exten- sions of perception verbs are grounded in interaction. Cognitive Linguistics 29, 371–406 (2018)

  8. [8]

    & Regier, T

    Kemp, C. & Regier, T. Kinship categories across languages reflect general communica- tive principles. Science 336, 1049–1054 (2012)

  9. [9]

    Thompson, B., Roberts, S. G. & Lupyan, G. Cultural influences on word meanings revealed through large-scale semantic alignment. Nature Human Behaviour 4, 1029– 1038 (2020). 14

  10. [10]

    Jackson, J. C. et al. Emotion semantics show both cultural variation and universal structure. Science 366, 1517–1522 (2019)

  11. [11]

    & List, J.-M

    Tjuka, A., Forkel, R. & List, J.-M. Universal and cultural factors shape body part vocabularies. Scientific Reports 14, 10486 (2024)

  12. [12]

    Thompson, B., Roberts, S. G. & Lupyan, G. Quantifying semantic similarity across languages in Annual Meeting of the Cognitive Science Society (2018), 2554–2559

  13. [13]

    Fedorenko, E., Piantadosi, S. T. & Gibson, E. A. Language is primarily a tool for communication rather than thought. Nature 630, 575–586 (2024)

  14. [14]

    & Centola, D

    Guilbeault, D., Baronchelli, A. & Centola, D. Experimental evidence for scale-induced category convergence across populations. Nature communications 12, 327 (2021)

  15. [15]

    Wang, X. & Bi, Y. Idiosyncratic tower of Babel: Individual differences in word-meaning representation increase as word abstractness increases. Psychological science 32, 1617– 1635 (2021)

  16. [16]

    & Lupyan, G

    Duan, Y. & Lupyan, G. Divergence in word meanings and its consequence for com- munication in Proceedings of the Annual Meeting of the Cognitive Science Society 45 (2023)

  17. [18]

    Barsalou, L. W. Grounded cognition: Past, present, and future. Topics in cognitive science 2, 716–724 (2010)

  18. [19]

    Explaining embodied cognition results

    Lakoff, G. Explaining embodied cognition results. Topics in cognitive science 4, 773–785 (2012)

  19. [20]

    Bergen, B. K. Louder than words: The new science of how the mind makes meaning (Basic Books, New York, 2012). 15

  20. [21]

    & Lupyan, G

    Lewis, M., Balamurugan, A., Zheng, B. & Lupyan, G. Characterizing variability in shared meaning through millions of sketches in Proceedings of the Annual Meeting of the Cognitive Science Society 43 (2021)

  21. [22]

    Malt, B. C. Representing the world in language and thought. Topics in Cognitive Science 16, 6–24 (2024)

  22. [23]

    Guilbeault, D. et al. Color associations in abstract semantic domains. Cognition 201, 104306 (2020)

  23. [24]

    Nadler, E. O. et al. Statistical or embodied? Comparing colorseeing, colorblind, painters, and Large Language Models in their processing of color metaphors. Cognitive Science 49, e70083 (2025)

  24. [25]

    Hand and Mind: What Gestures Reveal about Thought(University of Chicago Press, Chicago, 1992)

    McNeill, D. Hand and Mind: What Gestures Reveal about Thought(University of Chicago Press, Chicago, 1992)

  25. [26]

    J., Emmorey, K., Smith, J

    Xu, J., Gannon, P. J., Emmorey, K., Smith, J. F. & Braun, A. R. Symbolic gestures and spoken language are processed by a common neural system. Proceedings of the National Academy of Sciences 106, 20664–20669 (2009)

  26. [27]

    M., ¨Ozy¨ urek, A

    Willems, R. M., ¨Ozy¨ urek, A. & Hagoort, P. When language meets action: The neural integration of gesture and speech. Cerebral Cortex 17, 2322–2333 (2007)

  27. [28]

    M., Kita, S

    ¨Ozy¨ urek, A., Willems, R. M., Kita, S. & Hagoort, P. On-line integration of semantic in- formation from speech and gesture: Insights from event-related brain potentials.Journal of Cognitive Neuroscience 19, 605–616 (2007)

  28. [29]

    L., Humphries, C

    Fernandino, L., Tong, J.-Q., Conant, L. L., Humphries, C. J. & Binder, J. R. Decoding the information structure underlying the neural representation of concepts. Proceedings of the National Academy of Sciences 119, e2108091119 (2022)

  29. [30]

    Bechtold, L. et al. Brain signatures of embodied semantics and language: A consensus paper. Journal of cognition 6, 61 (2023). 16

  30. [31]

    Mukherjee, K. et al. Drawings of THINGS: A large-scale drawing dataset of 1,854 object concepts. Behavior Research Methods 58, 57 (2025)

  31. [32]

    N., Zhu, L

    Zhu, R., Kilonzo, T. N., Zhu, L. Z., Fan, J. E. & Frank, M. C. Cross-Contextual Vari- ability in Children’s Early Understanding of Visual Media. Topics in Cognitive Science (2025)

  32. [33]

    Long, B., Wang, Y., Christie, S., Frank, M. C. & Fan, J. E. Developmental changes in drawing production under different memory demands in a US and Chinese sample. Developmental Psychology 59, 1784 (2023)

  33. [34]

    E., Huey, H., Chai, Z

    Long, B., Fan, J. E., Huey, H., Chai, Z. & Frank, M. C. Parallel developmental changes in children’s production and recognition of line drawings of visual concepts. Nature Communications 15, 1191 (2024)

  34. [35]

    Yu, Q. et al. Sketch me that shoe in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2016), 799–807

  35. [36]

    Xu, P. et al. Deep learning for free-hand sketch: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 285–312 (2022)

  36. [37]

    Quick, Stat!: A Statistical Analysis of the Quick, Draw! Dataset

    Fernandez-Fernandez, R., Victores, J. G., Estevez, D. & Balaguer, C. Quick, stat!: A statistical analysis of the quick, draw! dataset. arXiv preprint arXiv:1907.06417 (2019)

  37. [38]

    A Neural Representation of Sketch Drawings

    Ha, D. & Eck, D. A neural representation of sketch drawings. arXiv preprint arXiv:1704.03477 (2017)

  38. [39]

    Xu, P. et al. Sketchmate: Deep hashing for million-scale human sketch retrieval in Pro- ceedings of the IEEE conference on computer vision and pattern recognition (2018), 8090–8098

  39. [40]

    Xu, P., Joshi, C. K. & Bresson, X. Multigraph transformer for free-hand sketch recog- nition. IEEE Transactions on Neural Networks and Learning Systems 33, 5150–5161 (2021). 17

  40. [41]

    Lamb, A., Ozair, S., Verma, V. & Ha, D. Sketchtransfer: A new dataset for exploring detail-invariance and the abstractions learned by deep networks in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (2020), 963–972

  41. [42]

    Murphy, G. L. Is there an exemplar theory of concepts? Psychonomic bulletin & review 23, 1035–1042 (2016)

  42. [43]

    Goldstein, E. B. Cognitive psychology: Connecting mind, research, and everyday expe- rience (Cengage learning Stamford, CT, 2015)

  43. [44]

    & Mervis, C

    Rosch, E. & Mervis, C. B. Family resemblances: Studies in the internal structure of categories. Cognitive psychology 7, 573–605 (1975)

  44. [45]

    Rogers, T. T. et al. Structure and deterioration of semantic memory: a neuropsycholog- ical and computational investigation. Psychological review 111, 205 (2004)

  45. [46]

    Medin, D. L. & Schaffer, M. M. Context theory of classification learning. Psychological review 85, 207 (1978)

  46. [47]

    Smith, E. E. & Medin, D. L. Categories and concepts (Harvard University Press, 1981)

  47. [48]

    Six views of embodied cognition

    Wilson, M. Six views of embodied cognition. Psychonomic bulletin & review 9, 625–636 (2002)

  48. [50]

    D., Guillaume, J.-L., Lambiotte, R

    Blondel, V. D., Guillaume, J.-L., Lambiotte, R. & Lefebvre, E. Fast unfolding of commu- nities in large networks. Journal of statistical mechanics: theory and experiment 2008, P10008 (2008)

  49. [52]

    J., Park, P

    Atari, M., Xue, M. J., Park, P. S., Blasi, D. E. & Henrich, J. Which Humans? 2023. 18

  50. [53]

    Michel, J.-B. et al. Quantitative analysis of culture using millions of digitized books. science 331, 176–182 (2011)

  51. [54]

    Watching a language model learning chess in Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021)(2021), 1369–1379

    St¨ ockl, A. Watching a language model learning chess in Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021)(2021), 1369–1379

  52. [55]

    & Hoyos-Idrobo, A

    Loyola, P., Marrese-Taylor, E. & Hoyos-Idrobo, A. Perceptual structure in the absence of grounding: the impact of abstractedness and subjectivity in color language for LLMs in Findings of the Association for Computational Linguistics: EMNLP 2023 (2023), 1536–1542

  53. [56]

    Piantadosi, S. T. et al. Why concepts are (probably) vectors. Trends in Cognitive Sci- ences 28, 844–856 (2024)

  54. [57]

    Frank, M. C. & Goodman, N. D. Cognitive modeling using artificial intelligence. Annual Review of Psychology 777:543-566 (2026)

  55. [58]

    & Griffiths, T

    Marjieh, R., Sucholutsky, I., van Rijn, P., Jacoby, N. & Griffiths, T. L. Large language models predict human sensory judgments across six modalities. Scientific Reports 14, 21445 (2024). Methods Data Processing To build the dataset, we keep only sketches that have been recognized by the neural network in the QuickDraw game. To prevent imbalance from over...

  56. [59]

    Edge similarity : the Jaccard index of edge sets, i.e., the proportion of shared edges relative to the union of edges in both networks

  57. [60]

    This captures local structural similarity even when the exact edges differ

    Neighborhood similarity: for each node, we compute the Jaccard index between its sets of neighbors in the two networks, and then average across all nodes. This captures local structural similarity even when the exact edges differ

  58. [61]

    This measure captures whether networks produce similar higher-level groupings of countries

    Community similarity: we compare Louvain communities detected in the two networks using the Normalized Mutual Information (NMI) score. This measure captures whether networks produce similar higher-level groupings of countries. We repeat the procedure by considering the two methods for network filtering (i.e., threshold on strongest edges and disparity fil...

  59. [62]

    Oquab, M. et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193 (2023)

  60. [63]

    UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

    McInnes, L., Healy, J. & Melville, J. Umap: Uniform manifold approximation and pro- jection for dimension reduction. arXiv preprint arXiv:1802.03426 (2018)

  61. [64]

    A density-based algorithm for dis- covering clusters in large spatial databases with noise in kdd 96 (1996), 226–231

    Ester, M., Kriegel, H.-P., Sander, J., Xu, X., et al. A density-based algorithm for dis- covering clusters in large spatial databases with noise in kdd 96 (1996), 226–231. 29

  62. [65]

    A., Campello, R

    Moulavi, D., Jaskowiak, P. A., Campello, R. J., Zimek, A. & Sander, J. Density-based clustering validation in Proceedings of the 2014 SIAM international conference on data mining (2014), 839–847

  63. [66]

    Brysbaert, M., Warriner, A. B. & Kuperman, V. Concreteness ratings for 40 thousand generally known English word lemmas. Behavior research methods 46, 904–911 (2014)

  64. [67]

    & Carney, J

    Lynott, D., Connell, L., Brysbaert, M., Brand, J. & Carney, J. The Lancaster Sensori- motor Norms: multidimensional measures of perceptual and action strength for 40,000 English words. Behavior research methods 52, 1271–1291 (2020)

  65. [68]

    Llama-3.3-70B-Instruct https://huggingface.co/meta- llama/Llama- 3.3- 70B-Instruct

    Meta. Llama-3.3-70B-Instruct https://huggingface.co/meta- llama/Llama- 3.3- 70B-Instruct. 2024

  66. [69]

    & Dean, J

    Mikolov, T., Chen, K., Corrado, G. & Dean, J. Word2Vec Google News 300-dimensional embeddings Pre-trained model, accessed 2025-03-01. https://huggingface.co/fse/ word2vec-google-news-300

  67. [70]

    GTE Multilingual Base Pre-trained BERT-based model, accessed 2025- 03-01

    Alibaba-NLP. GTE Multilingual Base Pre-trained BERT-based model, accessed 2025- 03-01. https://huggingface.co/Alibaba-NLP/gte-multilingual-base

  68. [71]

    Muthukrishna, M. et al. Beyond Western, Educated, Industrial, Rich, and Democratic (WEIRD) psychology: Measuring and mapping scales of cultural and psychological dis- tance. Psychological science 31, 678–701 (2020)

  69. [72]

    ´A., Bogun´ a, M

    Serrano, M. ´A., Bogun´ a, M. & Vespignani, A. Extracting the multiscale backbone of complex weighted networks. Proceedings of the national academy of sciences 106, 6483– 6488 (2009)

  70. [73]

    E., Strogatz, S

    Newman, M. E., Strogatz, S. H. & Watts, D. J. Random graphs with arbitrary degree distributions and their applications. Physical review E 64, 026118 (2001). 30 Acknowledgements A.P and L.M.A. acknowledge funding from Carlsberg Foundation Project COCOONS (Grant ID: CF21-0432). Author Contributions A.P., M.M., D.G., L.M.A, and A.B. designed the project. A.P...