Pith. sign in

REVIEW 3 major objections 6 minor 66 references

Billions of Sketches Reveal Hidden Cultural Variation in Human Concepts

T0 review · 3 major / 6 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read Common concepts unfold into multiple visual exemplars, and sketch-based similarity tracks cultural differences better than word-based similarity does.

desk verdict Impressive scale and honest analysis, but the headline 32% result is undercut by an unfair language baseline that collapses same-language countries. read the letter →

arxiv 2607.07267 v2 pith:H6I2WMPQ submitted 2026-07-08 cs.CY cs.CLphysics.soc-ph

classification cs.CYcs.CLphysics.soc-ph
keywords conceptualstructureculturalvariationsketchembeddingsvisualexemplarswordcross-culturalsimilarityembodiedcognitionmultimodalconcepts
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that how people draw common concepts — at massive scale — exposes cultural variation in mental concepts that word-based measures hide. Analyzing 2.6 billion drawings of 344 everyday concepts from 236 countries, it finds that most concepts split into a few stable visual exemplars rather than a single prototype, and that the geometry of sketch-based similarity diverges strongly from word-embedding similarity. The central quantitative claim is that country-to-country similarities computed from drawings align about 32% better with established cultural distances than similarities computed from translated word embeddings. If true, apparent universality of concepts is modality-dependent: words, as lossy communication devices, compress away variation that visual imagination preserves. A sympathetic reader would care because this offers a direct, non-linguistic, high-resolution probe of cultural conceptual diversity and challenges purely text-based models of human concepts.

What carries the argument

The central mechanism is the one-to-many mapping from a concept word to multiple visual exemplar clusters. Drawings are embedded in a visual similarity space, clustered into stable forms with noise separated, and each country is represented by its odds-ratio profile of cluster usage across concepts; these profiles yield a country similarity network that is compared with a language-based network built from translated concept names in multilingual word embeddings and with a cultural network built from survey-based value distances. The comparison uses edge, neighborhood, and community overlap measures relative to a null model. The load-bearing quantity is the 32% median ratio of image-culture s

What would settle it

A falsifying test is to recompute the country similarity network after matching or reweighting participants across countries on age, education, and internet access (or controlling for GDP per capita); if the image-culture alignment drops to the language-culture level, the cultural signal is an artifact of participation bias. Alternatively, a controlled drawing study with representative national samples could check whether the same 32% advantage appears when demographics are balanced.

Watch

Extended reading notes

Core claim

The discovery is that collective visual representations of common concepts are organized into multiple recurrent exemplars (median of two per concept; e.g., pizza drawn as a slice or a whole pie, fish facing left or right), that these visual geometries are nearly uncorrelated with word-based semantic geometries (rank correlation around 0.098), and that sketch-derived cross-national similarities match established cultural distances more closely than word-derived similarities do — a median improvement of 32% across network metrics and thresholds. The authors present this as evidence that conceptual universality depends on measurement modality: language compresses rich experiential variation in

Load-bearing premise

The central claim rests on the assumption that country-level drawing pools, though dominated by US and anglophone users and biased toward digitally privileged participants, still represent each country's culture well enough that sketch similarity tracks conceptual culture rather than shared participation demographics or development levels.

Editorial extensions

If this is right

  • Concepts that look universal in word-based analyses may show substantial variation when measured through drawing, so universality claims should specify the modality of measurement.
  • Text-only embedding models are likely to under-represent the cultural and embodied structure of human concepts, motivating multimodal training that includes visual or sensory data.
  • Large-scale sketch data can serve as a complementary tool for mapping cultural distances between countries, at least for the digitally connected populations represented in the data.
  • The visual-exemplar clustering of concepts provides a quantitative way to study within-concept cultural variability and its links to embodied experience.
  • Concepts strongly tied to hand/arm interaction are more visually coherent, suggesting embodied interaction shapes shared visual representations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the modality-dependence result holds, other non-linguistic modalities — emoji use, product images, gestures, sound — might be equally informative, and combining modalities could map cultural conceptual structure more completely than any single channel.
  • A direct testable extension is to re-weight or stratify the drawing sample by demographics (age, education, internet access) or to control for GDP and connectivity; if the image-culture alignment survives those controls, the cultural-signal interpretation is much stronger.
  • The 32% advantage may partly reflect that both drawings and the cultural survey come from people, whereas word embeddings come from text corpora with different population biases; aligning all measures to the same respondent population would sharpen the comparison.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper analyzes 2.6 billion QuickDraw sketches of 344 concepts from 236 countries. It first shows that sketches of a concept form multiple distinct visual exemplar clusters rather than a single prototype, and that clusterability correlates selectively with haptic and hand/arm sensorimotor properties. It then compares image-based concept embeddings with word embeddings, reporting low rank correspondence (macro-average 0.098) between visual and linguistic similarity rankings. Finally, it constructs country networks from sketch behavior and from primary-language word embeddings, and reports that the image-based country network aligns 32% more closely (median across conditions) with World Values Survey cultural distances than the language-based network does. The paper's central claim is that apparent conceptual universality is modality-dependent: visual representations preserve cultural variation that word embeddings compress away.

Significance. If the central comparison were fair, this would be an important contribution. The empirical scale is exceptional, and the robustness program is genuinely thorough: multiple embedding models, several edge thresholds, three node counts, two WVS matching criteria, WVS sub-dimensions, disparity filtering, partial Mantel tests, and code availability all strengthen the descriptive claims about sketch structure. The paper also makes a falsifiable prediction—image-based country similarities should track cultural distances better than word-based similarities—which is the right kind of claim to test. However, the language baseline is currently too coarse to support the headline 32% claim: it assigns each country one primary national language, so countries sharing a language receive identical profiles. The comparison therefore pits a rich behavioral measure against a drastically simplified linguistic label, and the main quantitative conclusion is at risk. The sampling bias acknowledged in the Discussion is also not controlled for, further threatening the cultural-inference component. With a fairer language baseline and demographic controls, the paper's thesis could be established; without the

major comments (3)
  1. [Methods, Networks; Results, Fig. 4; Table 3] The language-based network is built from each country's primary national language, so countries sharing a primary language (US, UK, Australia, Canada, India, Nigeria, Philippines for English) receive identical language profiles. The image-based network, by contrast, uses country-specific sketch behavior. The reported 32% image-over-language improvement and the partial Mantel result therefore compare a rich behavioral matrix against a country-level language label. This does not establish that 'words compress cultural variation'; it may only show that a single national-language label discards within-language cultural variation. The authors run extensive robustness checks on thresholds and node counts, but no robustness check on this baseline construction. A fairer baseline should use country-specific word usage, multilingual per-country distributions, or at least a language-family/area-lev
  2. [Discussion, Limitations; Methods, Data Processing; Table 1] The dataset is 41.3% US, the game interface is English, and the authors concede that participation is likely biased toward socioeconomically privileged cohorts (Discussion, Limitations). Country-level sketch similarity could therefore reflect shared participation demographics, internet penetration, or development gradients rather than conceptual structure. The alignment between image-based country networks and WVS cultural distances could be an artifact of the same demographic gradient driving both. The authors acknowledge the bias but do not control for it—there is no robustness check against GDP per capita, internet penetration, or English proficiency. At minimum, the paper should report partial correlations of the image-culture alignment with these variables, or show that the result survives when restricting to high-participation countries or to countries above a participation thresho
  3. [Methods, Clustering and Grid Components Clustering; SI Figs. 6, 9] The image-based network and all exemplar-cluster analyses depend on the cluster definitions, but these definitions are not independently validated. The grid-based high-density threshold is selected by maximizing precision against DBSCAN cluster labels on the same dataset (SI, Grid Components Clustering), and the clusterability noise threshold is the local minimum of the same data's noise distribution. This internal tuning risks overfitting the cluster solution to the specific clustering algorithm and dataset. The robustness checks cover embedding choice, edge thresholds, node counts, and filtering methods, but not the clustering pipeline itself. The authors should demonstrate cluster stability under subsampling and alternative clustering algorithms, or validate a sample of clusters against human judgments. This is less central than the language-baseline issue, but it is load-bearing for
minor comments (6)
  1. [SI Table 1] The country code for Latvia appears as 'L V' (with a space); it should be 'LV'.
  2. [Fig. 2 caption] The caption says 'The density distribution of property scores is reported on the x-axis,' but it is unclear what is being plotted. Please clarify whether this is a histogram, a density curve, or something else.
  3. [Data and Code Availability] The full dataset was shared under an NDA, and only a 50M sample is publicly available. Please state explicitly whether the released code can reproduce the main analyses on the public sample, and what exactly differs when using the full 2.6B dataset.
  4. [Word vs. Image Semantics] The macro-average 0.098 is described as a rank correlation across multiple metrics, but Rank-Biased Overlap, top-10 overlap, and Kendall's tau are not directly comparable. Please report each metric separately with confidence intervals, or clarify which single metric the 0.098 refers to.
  5. [SI, Clusterability robustness] The robustness of clusterability to embedding choice is tested on only ten sampled concepts (SI). The small sample should be acknowledged in the main text, and the 0.833 Spearman correlation should be accompanied by a confidence interval.
  6. [SI Fig. 13 caption] The caption says the word-based network is 'mapped to the coordinates of the image-based one,' but the word-based network is not initially embedded in a coordinate space. Please clarify how the coordinates were assigned.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: central claims are empirical comparisons against external benchmarks; minor self-citation and internal tuning do not force the results.

full rationale

The paper's load-bearing claims—multiple visual exemplars per concept, haptic-related clusterability, low image–word rank agreement (0.098), and the 32% higher alignment of sketch-based country networks with WVS cultural distances—are not defined into existence. Each is computed from sketch statistics and then compared with independent external benchmarks (Brysbaert/Lynott psycholinguistic norms, Word2Vec/BERT embeddings, and Cultural Fixation scores from WVS). No equation equates a fitted parameter with the reported outcome. The grid-clustering threshold is tuned against DBSCAN labels from the same dataset, but this is model selection for an auxiliary method, and the clusterability correlations use external concept-property ratings; it does not constitute a prediction forced by construction. The only self-citation appearing in a conceptual premise (ref. 13, Guilbeault–Baronchelli–Centola, 'words are a lossy medium') is not load-bearing: the paper's own empirical divergence measures carry the argument. The language-based country network uses only each country's primary national language, making same-language countries identical. This is a legitimate fairness/validity concern about the baseline, but the observed image–culture alignment is an external empirical outcome, not an equivalence-by-construction. The score reflects one minor non-load-bearing self-citation and an internal parameter-tuning step, not a circular derivation.

Assumptions & free parameters 8 free parameters · 7 assumptions · 1 invented entities

The paper is an empirical analysis, so its central claim rests mainly on domain assumptions about what the sketches measure (concepts vs. ambiguous English labels), who draws them (representativeness), and how countries and languages are mapped. No physical constants or established model parameters are pulled in; the fitted choices are clustering and network thresholds, all tuned on the same dataset, which means the descriptive statistics (median 2 clusters, clusterability, haptic correlation) are partly self-constituted by the pipeline.

free parameters (8)
  • DBSCAN epsilon (per concept) = not reported per category (selected via DBCV)
    Chosen by density-based cluster validity optimization on the same sketch data it is applied to; directly determines the number and size of visual clusters per concept, i.e., the paper's 'median 2 clusters' result (Methods, Clustering).
  • Minimum cluster size (1% of concept sketches) = 1%
    Hand-chosen cutoff below which clusters are reassigned to noise, shaping clusterability values and the haptic-correlation result (Methods, Clustering).
  • Grid high-density percentile = 60th percentile
    Selected by maximizing precision against DBSCAN labels on clusterable concepts from the same dataset; the grid pipeline defines exemplars for non-clusterable concepts (Methods, Grid Components Clustering; SI Fig. 9).
  • Clusterability noise threshold = local minimum of bimodal noise distribution
    Data-derived cutoff separating clusterable from non-clusterable concepts; clusterability is the input to the haptic-correlation result (Methods, Clustering; SI Fig. 6).
  • PCA dimensionality = 40
    Hand-chosen intermediate dimension between DINOv2 384-d embeddings and UMAP 2-d projection; affects cluster geometry (Methods, Data Processing).
  • Per-country sketch cap = 10,000 per category-country
    Hand-chosen cap to rebalance for over-represented countries such as the US; changes the odds-ratio profiles used in the country network (Methods, Data Processing).
  • UMAP hyperparameters (n_neighbors, min_dist) = not stated
    The clustering and all downstream cluster counts depend on the UMAP 2D embedding, but hyperparameters are not reported in the text.
  • Disparity-filter alpha_t = 0.25
    Chosen to retain 98% of the giant connected component in the culture network; used in the robustness version of the network comparison (Methods, Networks).
assumptions (7)
  • domain assumption A sketch produced in response to an English concept label directly reflects the concept's mental representation rather than ambiguity of the (English) stimulus
    The game prompt is the same English word for all players; variation across countries is interpreted as conceptual/cultural variation, but could partly reflect differential interpretation of the untranslated label (Introduction; Results).
  • domain assumption Users aggregated by IP-inferred country code represent that country's culture
    Country-level networks and the 32% cultural-alignment claim treat IP country as cultural membership; the authors note IP inference 'introduces additional uncertainty' (Methods, Networks; Discussion).
  • domain assumption The QuickDraw participant pool within each country is representative of that country's general population
    Load-bearing for any cross-national claim; the authors explicitly concede the sample is anglophone-, access-, and privilege-biased (Discussion, Limitations).
  • domain assumption Each country can be represented by a single primary national language for word-based similarity
    Countries sharing a primary language receive identical language profiles, collapsing all within-language cultural variation; stated in Methods, Networks.
  • domain assumption WVS-based Cultural Fixation distances are an appropriate external benchmark of cultural similarity for comparison with sketch and word networks
    'Established cultural distances' in the headline is operationalized by the Cultural Fixation index from the WVS 2010-2014 wave; different waves or survey dimensions change the benchmark (Methods, Networks; SI).
  • domain assumption Clusterability (fraction of sketches assigned to substantive clusters) measures how unambiguously a concept is visually represented, not other properties such as how easy the object is to draw
    This identification underpins the haptic-clusterability result (Results, Figure 2).
  • domain assumption Pretrained self-supervised visual embeddings (DINOv2, CLIP) provide a valid similarity space for human sketches
    All image-similarity results flow through these embeddings; robustness is checked for cluster counts on 9 sampled concepts, not for the network-level results (Methods; SI).
invented entities (1)
  • Visual exemplar attractor clusters
    purpose: Interpretive construct equating algorithmic clusters of sketches with distinct cognitive exemplars of a concept (e.g., pizza-slice versus whole-pizza)
    No falsifiable handle outside this paper; robustness to swapping DINOv2 for CLIP is internal evidence only, and the clusters are not validated against external behavioral predictions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Billions of Sketches Reveal Hidden Cultural Variation in Human Concepts." pith.science (2026). https://pith.science/paper/H6I2WMPQ

@misc{pith2026260707267,
  author       = {Pith},
  title        = {Pith review of: Billions of Sketches Reveal Hidden Cultural Variation in Human Concepts},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H6I2WMPQ}},
  note         = {Machine review of arXiv:2607.07267}
}
read the original abstract

Claims about the universality of human concepts have been predominantly assessed through linguistic similarity across languages and cultures. However, words are effective as communication devices because they compress rich experiential variation into shared conventions, potentially obscuring hidden individual and cultural differences in how concepts are mentally represented. Here, we analyse 2.6 billion human-made sketches of common concepts from 236 countries and territories to examine conceptual structure through people's visual imagination. Consistent with recent work on image-based cognition, we find that single concepts unfold into multiple distinct visual exemplars, revealing latent information about similarities and differences in conceptual structure across cultures. This variation is strongest for concepts involving haptic interaction, suggesting that visual imagery reflects variation in embodied experience as much as conventional definitions. Comparing embedding models of sketches with word embedding models across languages, we find that their geometries diverge, with visual representations preserving rich semantic and cultural structure that language models compress. Cross-cultural similarities derived from sketches align 32% more closely with established cultural distances than do text-based measures. Together, these results suggest that patterns of human conceptual universality may depend critically on the modality through which concepts are measured, with large-scale sketching providing a direct, high-resolution probe of conceptual diversity across embodied and cultural dimensions of thought.

Figures

Figures reproduced from arXiv: 2607.07267 by the authors.

Figure 1
Figure 1. Sketches of concepts provided by people around the world organize into distinct visual clusters. Visual clusters for six representative concepts (a–f) selected from the 344 available. Each point represents a drawing projected into a two-dimensional latent visual space. Points are colored according to their algorithmically assigned cluster, with gray points denoting drawings classified as random noise. Each cluster i… view at source ↗
Figure 2
Figure 2. Clusterability of sketches is selectively associated with haptic and sen￾sorimotor conceptual properties. Correlation between clusterability and conceptual properties of objects. Blue bars correspond to significant values (α = 0.05, after Bonfer￾roni correction), grey to non-significant ones. The density distribution of property scores is reported on the x -axis. To ensure that the correlations were not driven by sy… view at source ↗
Figure 3
Figure 3. Image- and word-based concept networks exhibit divergent large-scale patterns of inter-cultural distances and clustering. Networks of countries based on sketch similarity (left) and word similarity (right). The top 100 nodes by number of sketches are shown. Colors denote structural communities of countries that share a high level of similarity with one another within each network, as identified by the Louvain algori… view at source ↗
Figures from the paper (19 more)
Figure 4
Figure 4. Figure 4: Cultural similarity aligns more closely with image-based than word￾based concept networks. Comparison between image- and language-based network sim￾ilarity with the cultural network across different metrics, expressed as ratios relative to a baseline defined by a null …
Figure 5
Figure 5. Figure 5: Distribution of number of seconds passed before the sketched was recognized [PITH_FULL_IMAGE:figures/full_fig_p032_5.png]
Figure 5
Figure 5. Figure 5: Distribution of number of seconds passed before the sketched was recognized [PITH_FULL_IMAGE:figures/full_fig_p030_5.png]
Figure 6
Figure 6. Figure 6: Distributions of the percentage of noise in clustering (left) and the number of [PITH_FULL_IMAGE:figures/full_fig_p033_6.png]
Figure 7
Figure 7. Figure 7: Superposition of randomly sampled drawings within each cluster, for six represen [PITH_FULL_IMAGE:figures/full_fig_p034_7.png]
Figure 8
Figure 8. Figure 8: Steps of the grid-based pipeline on a selected example concepts ( [PITH_FULL_IMAGE:figures/full_fig_p034_8.png]
Figure 8
Figure 8. Figure 8: Steps of the grid-based pipeline on a selected example concepts ( [PITH_FULL_IMAGE:figures/full_fig_p035_8.png]
Figure 9
Figure 9. Figure 9: Performance of the alternative clustering methodology on clusterable concept cat [PITH_FULL_IMAGE:figures/full_fig_p035_9.png]
Figure 9
Figure 9. Figure 9: Performance of the alternative clustering methodology on clusterable concept cat [PITH_FULL_IMAGE:figures/full_fig_p036_9.png]
Figure 10
Figure 10. Figure 10: Similarity of rankings considering word and image embeddings, together with a [PITH_FULL_IMAGE:figures/full_fig_p036_10.png]
Figure 10
Figure 10. Figure 10: Communities of countries in the image embeddings-based network, plotted on a [PITH_FULL_IMAGE:figures/full_fig_p037_10.png]
Figure 11
Figure 11. Figure 11: NMI score for robustness of community structure across percentages of strongest [PITH_FULL_IMAGE:figures/full_fig_p037_11.png]
Figure 12
Figure 12. Figure 12: Comparison of network countries’ similarity based on image and words. Each [PITH_FULL_IMAGE:figures/full_fig_p040_12.png]
Figure 12
Figure 12. Figure 12: NMI score for robustness of community structure across percentages of strongest [PITH_FULL_IMAGE:figures/full_fig_p038_12.png]
Figure 13
Figure 13. Figure 13: Comparison between image- and language-based network similarity with the cul [PITH_FULL_IMAGE:figures/full_fig_p041_13.png]
Figure 13
Figure 13. Figure 13: Comparison of network countries’ similarity based on image and words. Each [PITH_FULL_IMAGE:figures/full_fig_p039_13.png]
Figure 14
Figure 14. Figure 14: Comparison between image- and language-based network similarity with the [PITH_FULL_IMAGE:figures/full_fig_p040_14.png]
Figure 15
Figure 15. Figure 15: Comparison between image- and language-based network similarity with the [PITH_FULL_IMAGE:figures/full_fig_p041_15.png]
Figure 16
Figure 16. Figure 16: Comparison between image- and language-based network similarity with the [PITH_FULL_IMAGE:figures/full_fig_p043_16.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

66 extracted references · 4 linked inside Pith

  1. [1]

    & Kay, P

    Berlin, B. & Kay, P. Basic color terms: Their universality and evolution (Univ of Cali- fornia Press, 1969)

  2. [2]

    & Regier, T

    Kemp, C., Xu, Y. & Regier, T. Semantic typology and efficient communication. Annual Review of Linguistics 4, 109–128 (2018)

  3. [3]

    & Evans, J

    Lewis, M., Cahill, A., Madnani, N. & Evans, J. Local similarity and global variabil- ity characterize the semantic space of human languages. Proceedings of the National Academy of Sciences 120, e2300986120 (2023)

  4. [4]

    C., Jiang, S

    Xu, Y., Duong, K., Malt, B. C., Jiang, S. & Srinivasan, M. Conceptual relations predict colexification across languages. Cognition 201, 104280 (2020)

  5. [5]

    & Ran, Q

    Liang, Y., Xu, K. & Ran, Q. Shared structure of fundamental human experience revealed by polysemy network of basic vocabularies across languages. Scientific Reports 14, 5877 (2024)

  6. [6]

    & Tishby, N

    Zaslavsky, N., Kemp, C., Regier, T. & Tishby, N. Efficient compression in color naming and its evolution. Proceedings of the National Academy of Sciences 115, 7937–7942 (2018)

  7. [7]

    H., Norcliffe, E

    San Roque, L., Kendrick, K. H., Norcliffe, E. & Majid, A. Universal meaning exten- sions of perception verbs are grounded in interaction. Cognitive Linguistics 29, 371–406 (2018)

  8. [8]

    & Regier, T

    Kemp, C. & Regier, T. Kinship categories across languages reflect general communica- tive principles. Science 336, 1049–1054 (2012)

Show all 66 references
  1. [9]

    Thompson, B., Roberts, S. G. & Lupyan, G. Cultural influences on word meanings revealed through large-scale semantic alignment. Nature Human Behaviour 4, 1029– 1038 (2020). 23

  2. [10]

    Jackson, J. C. et al. Emotion semantics show both cultural variation and universal structure. Science 366, 1517–1522 (2019)

  3. [11]

    & List, J.-M

    Tjuka, A., Forkel, R. & List, J.-M. Universal and cultural factors shape body part vocabularies. Scientific Reports 14, 10486 (2024)

  4. [12]

    Thompson, B., Roberts, S. G. & Lupyan, G. Quantifying semantic similarity across languages in Annual Meeting of the Cognitive Science Society (2018), 2554–2559

  5. [13]

    & Centola, D

    Guilbeault, D., Baronchelli, A. & Centola, D. Experimental evidence for scale-induced category convergence across populations. Nature communications 12, 327 (2021)

  6. [14]

    Barsalou, L. W. Grounded cognition: Past, present, and future. Topics in cognitive science 2, 716–724 (2010)

  7. [15]

    Explaining embodied cognition results

    Lakoff, G. Explaining embodied cognition results. Topics in cognitive science 4, 773–785 (2012)

  8. [16]

    Bergen, B. K. Louder than words: The new science of how the mind makes meaning (Basic Books, New York, 2012)

  9. [17]

    & Lupyan, G

    Lewis, M., Balamurugan, A., Zheng, B. & Lupyan, G. Characterizing variability in shared meaning through millions of sketches in Proceedings of the Annual Meeting of the Cognitive Science Society 43 (2021)

  10. [18]

    Fedorenko, E., Piantadosi, S. T. & Gibson, E. A. Language is primarily a tool for communication rather than thought. Nature 630, 575–586 (2024)

  11. [19]

    Malt, B. C. Representing the world in language and thought. Topics in Cognitive Science 16, 6–24 (2024)

  12. [20]

    Guilbeault, D. et al. Color associations in abstract semantic domains. Cognition 201, 104306 (2020). 24

  13. [21]

    Nadler, E. O. et al. Statistical or embodied? Comparing colorseeing, colorblind, painters, and Large Language Models in their processing of color metaphors. Cognitive Science 49, e70083 (2025)

  14. [22]

    Hand and Mind: What Gestures Reveal about Thought (University of Chicago Press, Chicago, 1992)

    McNeill, D. Hand and Mind: What Gestures Reveal about Thought (University of Chicago Press, Chicago, 1992)

  15. [23]

    J., Emmorey, K., Smith, J

    Xu, J., Gannon, P. J., Emmorey, K., Smith, J. F. & Braun, A. R. Symbolic gestures and spoken language are processed by a common neural system. Proceedings of the National Academy of Sciences 106, 20664–20669 (2009)

  16. [24]

    M., ¨Ozy¨ urek, A

    Willems, R. M., ¨Ozy¨ urek, A. & Hagoort, P. When language meets action: The neural integration of gesture and speech. Cerebral Cortex 17, 2322–2333 (2007)

  17. [25]

    M., Kita, S

    ¨Ozy¨ urek, A., Willems, R. M., Kita, S. & Hagoort, P. On-line integration of semantic in- formation from speech and gesture: Insights from event-related brain potentials.Journal of Cognitive Neuroscience 19, 605–616 (2007)

  18. [26]

    L., Humphries, C

    Fernandino, L., Tong, J.-Q., Conant, L. L., Humphries, C. J. & Binder, J. R. Decoding the information structure underlying the neural representation of concepts. Proceedings of the National Academy of Sciences 119, e2108091119 (2022)

  19. [27]

    Bechtold, L. et al. Brain signatures of embodied semantics and language: A consensus paper. Journal of cognition 6, 61 (2023)

  20. [28]

    Mukherjee, K. et al. Drawings of THINGS: A large-scale drawing dataset of 1,854 object concepts. Behavior Research Methods 58, 57 (2025)

  21. [29]

    N., Zhu, L

    Zhu, R., Kilonzo, T. N., Zhu, L. Z., Fan, J. E. & Frank, M. C. Cross-Contextual Vari- ability in Children’s Early Understanding of Visual Media. Topics in Cognitive Science (2025)

  22. [30]

    Long, B., Wang, Y., Christie, S., Frank, M. C. & Fan, J. E. Developmental changes in drawing production under different memory demands in a US and Chinese sample. Developmental Psychology 59, 1784 (2023). 25

  23. [31]

    E., Huey, H., Chai, Z

    Long, B., Fan, J. E., Huey, H., Chai, Z. & Frank, M. C. Parallel developmental changes in children’s production and recognition of line drawings of visual concepts. Nature Communications 15, 1191 (2024)

  24. [32]

    Yu, Q. et al. Sketch me that shoe in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2016), 799–807

  25. [33]

    Xu, P. et al. Deep learning for free-hand sketch: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 285–312 (2022)

  26. [34]

    G., Estevez, D

    Fernandez-Fernandez, R., Victores, J. G., Estevez, D. & Balaguer, C. Quick, stat!: A statistical analysis of the quick, draw! dataset. arXiv preprint arXiv:1907.06417 (2019)

  27. [35]

    & Eck, D

    Ha, D. & Eck, D. A neural representation of sketch drawings. arXiv preprint arXiv:1704.03477 (2017)

  28. [36]

    Xu, P. et al. Sketchmate: Deep hashing for million-scale human sketch retrieval in Pro- ceedings of the IEEE conference on computer vision and pattern recognition (2018), 8090–8098

  29. [37]

    Xu, P., Joshi, C. K. & Bresson, X. Multigraph transformer for free-hand sketch recog- nition. IEEE Transactions on Neural Networks and Learning Systems 33, 5150–5161 (2021)

  30. [38]

    Lamb, A., Ozair, S., Verma, V. & Ha, D. Sketchtransfer: A new dataset for exploring detail-invariance and the abstractions learned by deep networks in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (2020), 963–972

  31. [39]

    Murphy, G. L. Is there an exemplar theory of concepts? Psychonomic bulletin & review 23, 1035–1042 (2016)

  32. [40]

    Goldstein, E. B. Cognitive psychology: Connecting mind, research, and everyday expe- rience (Cengage learning Stamford, CT, 2015). 26

  33. [41]

    & Mervis, C

    Rosch, E. & Mervis, C. B. Family resemblances: Studies in the internal structure of categories. Cognitive psychology 7, 573–605 (1975)

  34. [42]

    Rogers, T. T. et al. Structure and deterioration of semantic memory: a neuropsycholog- ical and computational investigation. Psychological review 111, 205 (2004)

  35. [43]

    Medin, D. L. & Schaffer, M. M. Context theory of classification learning. Psychological review 85, 207 (1978)

  36. [44]

    Smith, E. E. & Medin, D. L. Categories and concepts (Harvard University Press, 1981)

  37. [45]

    Six views of embodied cognition

    Wilson, M. Six views of embodied cognition. Psychonomic bulletin & review 9, 625–636 (2002)

  38. [46]

    Brysbaert, M., Warriner, A. B. & Kuperman, V. Concreteness ratings for 40 thousand generally known English word lemmas. Behavior research methods 46, 904–911 (2014)

  39. [47]

    & Carney, J

    Lynott, D., Connell, L., Brysbaert, M., Brand, J. & Carney, J. The Lancaster Sensori- motor Norms: multidimensional measures of perceptual and action strength for 40,000 English words. Behavior research methods 52, 1271–1291 (2020)

  40. [48]

    D., Guillaume, J.-L., Lambiotte, R

    Blondel, V. D., Guillaume, J.-L., Lambiotte, R. & Lefebvre, E. Fast unfolding of commu- nities in large networks. Journal of statistical mechanics: theory and experiment 2008, P10008 (2008)

  41. [49]

    A., Waltman, L

    Traag, V. A., Waltman, L. & Van Eck, N. J. From Louvain to Leiden: guaranteeing well-connected communities. Scientific reports 9, 5233 (2019)

  42. [50]

    Muthukrishna, M. et al. Beyond Western, Educated, Industrial, Rich, and Democratic (WEIRD) psychology: Measuring and mapping scales of cultural and psychological dis- tance. Psychological science 31, 678–701 (2020)

  43. [51]

    J., Park, P

    Atari, M., Xue, M. J., Park, P. S., Blasi, D. E. & Henrich, J. Which Humans? 2023

  44. [52]

    Michel, J.-B. et al. Quantitative analysis of culture using millions of digitized books. science 331, 176–182 (2011). 27

  45. [53]

    Watching a language model learning chess in Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021) (2021), 1369–1379

    St¨ ockl, A. Watching a language model learning chess in Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021) (2021), 1369–1379

  46. [54]

    & Hoyos-Idrobo, A

    Loyola, P., Marrese-Taylor, E. & Hoyos-Idrobo, A. Perceptual structure in the absence of grounding: the impact of abstractedness and subjectivity in color language for LLMs in Findings of the Association for Computational Linguistics: EMNLP 2023 (2023), 1536–1542

  47. [55]

    Piantadosi, S. T. et al. Why concepts are (probably) vectors. Trends in Cognitive Sci- ences 28, 844–856 (2024)

  48. [56]

    Frank, M. C. & Goodman, N. D. Cognitive modeling using artificial intelligence. Annual Review of Psychology 777:543-566 (2026)

  49. [57]

    & Griffiths, T

    Marjieh, R., Sucholutsky, I., van Rijn, P., Jacoby, N. & Griffiths, T. L. Large language models predict human sensory judgments across six modalities. Scientific Reports 14, 21445 (2024)

  50. [58]

    Oquab, M. et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193 (2023)

  51. [59]

    Radford, A. et al. Learning transferable visual models from natural language supervision in International conference on machine learning (2021), 8748–8763

  52. [60]

    & Melville, J

    McInnes, L., Healy, J. & Melville, J. Umap: Uniform manifold approximation and pro- jection for dimension reduction. arXiv preprint arXiv:1802.03426 (2018)

  53. [61]

    A density-based algorithm for dis- covering clusters in large spatial databases with noise in kdd 96 (1996), 226–231

    Ester, M., Kriegel, H.-P., Sander, J., Xu, X., et al. A density-based algorithm for dis- covering clusters in large spatial databases with noise in kdd 96 (1996), 226–231

  54. [62]

    A., Campello, R

    Moulavi, D., Jaskowiak, P. A., Campello, R. J., Zimek, A. & Sander, J. Density-based clustering validation in Proceedings of the 2014 SIAM international conference on data mining (2014), 839–847. 28

  55. [63]

    & Dean, J

    Mikolov, T., Chen, K., Corrado, G. & Dean, J. Word2Vec Google News 300-dimensional embeddings Pre-trained model, accessed 2025-03-01. https://huggingface.co/fse/ word2vec-google-news-300

  56. [64]

    GTE Multilingual Base Pre-trained BERT-based model, accessed 2025- 03-01

    Alibaba-NLP. GTE Multilingual Base Pre-trained BERT-based model, accessed 2025- 03-01. https://huggingface.co/Alibaba-NLP/gte-multilingual-base

  57. [65]

    ´A., Bogun´ a, M

    Serrano, M. ´A., Bogun´ a, M. & Vespignani, A. Extracting the multiscale backbone of complex weighted networks. Proceedings of the national academy of sciences 106, 6483– 6488 (2009)

  58. [66]

    E., Strogatz, S

    Newman, M. E., Strogatz, S. H. & Watts, D. J. Random graphs with arbitrary degree distributions and their applications. Physical review E 64, 026118 (2001). Acknowledgements A.P. and L.M.A. acknowledge funding from Carlsberg Foundation Project COCOONS (Grant ID: CF21-0432). Au...

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.