REVIEW 2 major objections 5 minor 36 references
Speech language models judge spoken pseudowords as round or pointed far less like humans than they judge shapes, missing the acoustic cues that drive human sound symbolism.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 13:48 UTC pith:N2IB4EKZ
load-bearing objection First real-speech bouba/kiki probe of SLMs; open models miss spectral tilt and fail crossmodal matching while vision is near ceiling—solid empirical localization, ordinary sampling caveats. the 2 major comments →
Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Speech language models' auditory judgments of spoken pseudowords align only weakly with human sound-symbolism ratings (best first-order Pearson r = 0.470 against a human ceiling of 0.843) and do not exploit the acoustic cues, particularly spectral tilt, that drive human intuitions; open-weight omni models cannot reliably match a heard sound to the corresponding shape, while a visual-only control yields near-ceiling shape ratings, localizing the failure to speech representations rather than shape perception.
What carries the argument
A four-experiment decomposition of the bouba/kiki effect into auditory forced-choice and graded-rating tasks, crossmodal sound-to-shape matching, and a visual-only shape-rating control, scored against human ceilings and second-order representational similarity analysis of acoustic-cue matrices.
Load-bearing premise
That low model-human correlations on isolated graded ratings of fixed human speech recordings, extracted by an LLM judge, directly measure a deficit in acoustic representation rather than prompt, sampling, or response-format artifacts.
What would settle it
Retrain or fine-tune an open-weight omni model so that its second-order RSA correlation with human spectral-tilt matrices approaches the human value of 0.431, then re-run the graded-rating and crossmodal matching experiments; if first-order agreement and matching rates remain near chance, the localization claim is false.
If this is right
- Perceptual alignment of speech models will not be solved by stronger vision towers alone.
- Training objectives that ignore low-level spectral and temporal cues will continue to produce outputs that feel incongruent to human listeners.
- Benchmarks that evaluate only text or synthetic audio will miss the speech-specific gap documented here.
- Layer-wise analyses can be used to test whether future audio encoders form the rounded/pointed distinction earlier than the final layers observed for current models.
Where Pith is reading between the lines
- The same speech-representation gap is likely to appear in any generative task that maps spoken names or onomatopoeia onto visual form, not only forced-choice matching.
- Contrastive or self-supervised pretraining that explicitly reconstructs spectral tilt and envelope features could raise human-model RSA without new labeled data.
- The late-layer emergence of decisions under the logit lens suggests that intermediate audio embeddings remain largely uninformative about sound symbolism, offering a diagnostic for future encoders.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks whether Speech Language Models (SLMs) exhibit the human bouba/kiki sound-symbolism effect when given real speech rather than text or images. Using human-recorded pseudowords and shapes from established psycholinguistic studies (Ćwiek et al., McCormick et al., Lacey et al.), it runs four experiments: forced-choice and graded auditory ratings (Exps. 1–2), crossmodal sound-to-shape matching (Exp. 3), and a visual-only shape-rating ablation (Exp. 4). Eight models (four audio-only, four omni) are tested with multi-language prompts, human ceilings, Mantel-tested RSA, acoustic-cue RSA, and a logit-lens analysis of Qwen3-Omni. The central claim is that auditory judgments align poorly with humans (best first-order Pearson r = 0.470 vs. ceiling 0.843) and miss spectral tilt and related cues, while open-weight omni models fail at sound–shape matching; near-ceiling visual ratings localize the deficit to speech representations rather than shape perception.
Significance. If the result holds, it supplies the first systematic, speech-grounded evaluation of sound symbolism in SLMs and a clean dissociation between intact visual shape perception and weak auditory/crossmodal alignment. The design is careful: independent human baselines, human ceilings, counterbalanced option order, multi-language prompts, Mantel RSA, acoustic-cue profiles (Table III), visual ablation (Table IV), and manual verification of the LLM judge. These strengths make the localization claim (failure on the auditory side) a useful, falsifiable target for future audio representation work and for human–machine perceptual alignment in creative SLM applications. The open-weight vs. frontier contrast and the late-layer logit-lens results further sharpen where the gap sits.
major comments (2)
- §IV-B and Table II: The central auditory-alignment claim rests on graded ratings at temperature 0.7 with 30 samples per item and free-form answers parsed by an LLM judge. Although the paper documents order/response biases in Exp. 1 and verifies the judge on 500 samples, it does not report a temperature-0 (or low-temperature) control or an alternative extraction protocol for Exp. 2. Because the best model reaches only r = 0.470 against a human ceiling of 0.843, residual sampling/prompt artifacts remain a load-bearing alternative explanation for the low correlations and the acoustic-cue RSA mismatch in Table III. A short ablation (temperature 0 or constrained decoding on a subset of the 536 items) would substantially strengthen the claim that the deficit is representational rather than procedural.
- §IV-C and Table IV (Exp. 3): Gemini3.5-Flash reaches perfect matching while showing only weak auditory alignment (r = 0.392) and near-zero acoustic-cue RSA. The paper’s plausible lexical/semantic route is noted but not tested (e.g., by replacing bouba/kiki with less familiar or scrambled pseudowords, or by probing whether the model can name the shapes without audio). Without that control, the open-weight failure is well supported, but the stronger claim that success requires acoustic rather than lexical routes remains under-specified for the frontier model.
minor comments (5)
- §III-A / Appendix B: The forced-choice prompt uses “spiky” while the graded prompts and human literature use “pointed”; a brief note on synonym choice would avoid any reader concern about label mismatch.
- Table I: Several models show extreme order or response biases (e.g., MiniCPM-o-4.5, Qwen3-Omni, Gemma4-E4B). The text already flags these; adding a short column or footnote with first-option preference rates would make the artifact transparent in the table itself.
- Fig. 2–3: The logit-lens figures are informative but dense; labeling the approximate layer index where max digit probability rises (and where band separation appears for images) would help readers without re-reading the caption.
- §II-B: The discussion of prior audio-domain work [18] is brief; one extra sentence on how synthetic labels and dictionary words differ from the present human-recorded pseudoword design would sharpen the novelty claim.
- Appendix A / C: Licences and ethical notes are thorough; ensuring the code release includes the exact prompt translations and judge templates (as promised) will aid reproducibility.
Circularity Check
No circularity: empirical model-vs-human correlations and RSA against independent prior psycholinguistic data; no fitted parameters re-predicted, no self-definitional steps, no load-bearing self-citations.
full rationale
The paper's central claims rest on querying SLMs with fixed human speech recordings (Ćwiek et al., McCormick et al.) and shapes (Lacey et al.), extracting free-form answers via an LLM judge, then computing first-order Pearson/Spearman correlations and second-order Mantel RSA against the same independent human rating matrices and acoustic-parameter matrices published by those prior groups. Human ceilings (split-half reliability) and acoustic-cue profiles (spectral tilt ρ=.431 for humans) are taken as external benchmarks; model scores are never used to define or fit the quantities later reported as 'alignment' or 'cue sensitivity.' Visual ablation (Exp. 4) and logit-lens layer trajectories are post-hoc diagnostics of the same scores, not constructions that force the auditory shortfall. No uniqueness theorems, ansätze, or self-citations by the present authors underwrite any load-bearing step. The derivation chain is therefore an ordinary empirical comparison, fully self-contained against external data.
Axiom & Free-Parameter Ledger
free parameters (3)
- sampling temperature =
0.7
- number of samples per item =
30
- unparseable-rate language filter threshold =
<10%
axioms (5)
- domain assumption Published human bouba/kiki and shape ratings (Ćwiek, McCormick, Lacey) are valid ground-truth measures of human sound symbolism.
- domain assumption Forced-choice and 7-point graded rating prompts elicit the same perceptual mapping that humans exhibit in the classic matching task.
- domain assumption LLM-as-judge extraction of free-form answers is accurate (verified on 500 samples).
- standard math Mantel permutation tests and Spearman–Brown reliability ceilings are appropriate for comparing dissimilarity matrices and human ceilings.
- domain assumption Spectral tilt and the other nine acoustic parameters of Lacey et al. are the relevant cues for human sound symbolism.
read the original abstract
Sound symbolism, the human tendency to map speech sounds to perceptual qualities such as roundness or sharpness, arises primarily from the acoustics of speech rather than spelling. Whether Speech Language Models (SLMs) share this tendency remains open, as prior evaluations rely on text or images rather than real speech. We study it using genuine human speech recordings, comparing model judgments against human data across the auditory, crossmodal, and visual components of the effect. We find that SLMs' auditory judgments align poorly with human perception and miss the acoustic cues, such as spectral tilt, that drive human intuitions, and open-weight models cannot reliably link a heard sound to its corresponding shape. With a visual-only control ruling out shape perception, the weakness localizes to how speech is represented, suggesting that perceptual alignment depends not on stronger vision but on speech representations that capture the cues humans hear.
Figures
Reference graph
Works this paper leans on
-
[1]
Hinton, J
L. Hinton, J. Nichols, and J. J. Ohala,Sound symbolism. Cambridge University Press, 1994
1994
-
[2]
Arbitrariness, iconicity, and systematicity in language,
M. Dingemanse, D. E. Blasi, G. Lupyan, M. H. Christiansen, and P. Monaghan, “Arbitrariness, iconicity, and systematicity in language,” Trends in cognitive sciences, vol. 19, no. 10, pp. 603–615, 2015
2015
-
[3]
Gestalt psychology,
W. K ¨ohler, “Gestalt psychology,”Psychologische forschung, vol. 31, no. 1, pp. XVIII–XXX, 1967
1967
-
[4]
Synaesthesia — a window into perception, thought and language,
V . S. Ramachandran and E. M. Hubbard, “Synaesthesia — a window into perception, thought and language,”Journal of Consciousness Studies, vol. 8, no. 12, pp. 3–34, 2001
2001
-
[5]
The bouba/kiki effect is robust across cultures and writing systems,
A. ´Cwiek, S. Fuchs, C. Draxler, E. L. Asu, D. Dediu, K. Hiovain, S. Kawahara, S. Koutalidis, M. Krifka, P. Lippuset al., “The bouba/kiki effect is robust across cultures and writing systems,”Philosophical Transactions of the Royal Society B: Biological Sciences, vol. 377, no. 1841, p. 20200390, 2021
2021
-
[6]
An ethological perspective on common cross-language utilization of F0 of voice,
J. J. Ohala, “An ethological perspective on common cross-language utilization of F0 of voice,”Phonetica, vol. 41, pp. 1–16, 1984
1984
-
[7]
A study in phonetic symbolism
E. Sapir, “A study in phonetic symbolism.”Journal of experimental psychology, vol. 12, no. 3, p. 225, 1929
1929
-
[8]
Stimulus parameters underlying sound-symbolic mapping of auditory pseudowords to visual shapes,
S. Lacey, Y . Jamal, S. M. List, K. McCormick, K. Sathian, and L. C. Nygaard, “Stimulus parameters underlying sound-symbolic mapping of auditory pseudowords to visual shapes,”Cognitive Science, vol. 44, no. 9, p. e12883, 2020
2020
-
[9]
Generating images from spoken descriptions,
X. Wang, T. Qiao, J. Zhu, A. Hanjalic, and O. Scharenborg, “Generating images from spoken descriptions,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 850–865, 2021
2021
-
[10]
Audio flamingo 3: Advancing audio intelligence with fully open large audio language models,
S. Ghosh, A. Goel, J. Kim, S. Kumar, Z. Kong, S.-g. Lee, C.-H. Yang, R. Duraiswami, D. Manocha, R. Valle, and B. Catanzaro, “Audio flamingo 3: Advancing audio intelligence with fully open large audio language models,” inAdvances in Neural Information Processing Systems, D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen, Eds.,...
2025
-
[11]
Step-audio 2 technical report,
B. Wu, C. Yan, C. Hu, C. Yi, C. Feng, F. Tian, F. Shen, G. Yu, H. Zhang, J. Liet al., “Step-audio 2 technical report,”arXiv preprint arXiv:2507.16632, 2025
Pith/arXiv arXiv 2025
-
[12]
D. Ding, Z. Ju, Y . Leng, S. Liu, T. Liu, Z. Shang, K. Shen, W. Song, X. Tan, H. Tanget al., “Kimi-audio technical report,”arXiv preprint arXiv:2504.18425, 2025
Pith/arXiv arXiv 2025
-
[13]
Minicpm-o 4.5: Towards real-time full-duplex omni-modal interaction,
J. Cui, B. Xu, C. Wang, T. Yu, W. Sun, Y . Xu, T. Wang, Z. He, W. Ma, T. Caiet al., “Minicpm-o 4.5: Towards real-time full-duplex omni-modal interaction,”arXiv preprint arXiv:2604.27393, 2026
Pith/arXiv arXiv 2026
-
[14]
Towards holistic evaluation of large audio-language models: A comprehensive survey,
C.-K. Yang, N. S. Ho, and H.-y. Lee, “Towards holistic evaluation of large audio-language models: A comprehensive survey,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V . Peng, Eds. Suzhou, China: Association for Computational Linguistics, Nov. 2025, pp. 10 1...
2025
-
[15]
J. Xu, Z. Guo, H. Hu, Y . Chu, X. Wang, J. He, Y . Wang, X. Shi, T. He, X. Zhuet al., “Qwen3-omni technical report,”arXiv preprint arXiv:2509.17765, 2025
Pith/arXiv arXiv 2025
-
[16]
Kiki or bouba? sound symbolism in vision-and-language models,
M. Alper and H. Averbuch-Elor, “Kiki or bouba? sound symbolism in vision-and-language models,”Advances in Neural Information Process- ing Systems, vol. 36, pp. 78 347–78 359, 2023
2023
-
[17]
Evidence for systematic semantic structure in individual phonemes,
G. Zhao, “Evidence for systematic semantic structure in individual phonemes,”arXiv preprint arXiv:2603.17306, 2026
Pith/arXiv arXiv 2026
-
[18]
Do language models associate sound with meaning? a multimodal study of sound symbolism,
J. Jeong, S. Lee, J. Lee, S. Han, and Y . Yu, “Do language models associate sound with meaning? a multimodal study of sound symbolism,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 37, 2026, pp. 31 247–31 255
2026
-
[19]
Interpreting GPT: the logit lens,
nostalgebraist, “Interpreting GPT: the logit lens,” https://www.lesswron g.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens, August 2020
2020
-
[20]
Dynamic-SUPERB phase-2: A collaboratively expanding benchmark for measuring the capabilities of spoken language models with 180 tasks,
C.-y. Huanget al., “Dynamic-SUPERB phase-2: A collaboratively expanding benchmark for measuring the capabilities of spoken language models with 180 tasks,” inThe Thirteenth International Conference on Learning Representations, 2025
2025
-
[21]
V oicebench: Benchmarking llm-based voice assistants,
Y . Chen, X. Yue, C. Zhang, X. Gao, R. T. Tan, and H. Li, “V oicebench: Benchmarking llm-based voice assistants,”Transactions of the Associa- tion for Computational Linguistics, vol. 14, pp. 378–398, 2026
2026
-
[22]
Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Mod- els,
K.-H. Luet al., “Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Mod- els,” inInterspeech 2025, 2025, pp. 2078–2082
2025
-
[23]
Listen, think, and understand,
Y . Gong, H. Luo, A. Liu, L. Karlinsky, and J. R. Glass, “Listen, think, and understand,” inInternational Conference on Learning Representations, B. Kim, Y . Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y . Sun, Eds., vol. 2024, 2024, pp. 18 516–18 545. [Online]. Available: https://proceedings.iclr.cc/paper files/paper/2024/ file/510d0935b543a29d686f93...
2024
-
[24]
AudioBench: A universal benchmark for audio large language models,
B. Wang, X. Zou, G. Lin, S. Sun, Z. Liu, W. Zhang, Z. Liu, A. Aw, and N. F. Chen, “AudioBench: A universal benchmark for audio large language models,” inProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and...
2025
-
[25]
Mmau: A massive multi-task audio understanding and reasoning benchmark,
S. Sakshi, U. Tyagi, S. Kumar, A. Seth, R. Selvakumar, O. Nieto, R. Duraiswami, S. Ghosh, and D. Manocha, “Mmau: A massive multi-task audio understanding and reasoning benchmark,” 2024. [Online]. Available: https://arxiv.org/abs/2410.19168
Pith/arXiv arXiv 2024
-
[26]
S. Kumar, ˇS. Sedl ´aˇcek, V . Lokegaonkar, F. L ´opez, W. Yu, N. Anand, H. Ryu, L. Chen, M. Pli ˇcka, M. Hlav ´aˇceket al., “Mmau-pro: A challenging and comprehensive benchmark for holistic evaluation of audio general intelligence,”arXiv preprint arXiv:2508.13992, 2025
Pith/arXiv arXiv 2025
-
[27]
Dynamic-superb: To- wards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech,
C.-Y . Huang, K.-H. Lu, S.-H. Wang, C.-Y . Hsiao, C.-Y . Kuan, H. Wu, S. Arora, K.-W. Chang, J. Shi, Y . Peng, R. Sharma, S. Watanabe, B. Ramakrishnan, S. Shehata, and H.-Y . Lee, “Dynamic-superb: To- wards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech,” inICASSP 2024 - 2024 IEEE International Conference on Acoustics,...
2024
-
[28]
Echomind: An interrelated multi-level benchmark for evaluating empathetic speech language models,
L. Zhou, L. Yu, Y . Lyu, Y . Lin, Z. Zhao, J. Ao, Y . Zhang, B. Wang, and H. Li, “Echomind: An interrelated multi-level benchmark for evaluating empathetic speech language models,”arXiv preprint arXiv:2510.22758, 2025
arXiv 2025
-
[29]
Investigating iconicity in vision-and-language models: A case study of the bouba/kiki effect in japanese models,
H. Iida and H. Funakura, “Investigating iconicity in vision-and-language models: A case study of the bouba/kiki effect in japanese models,” in Proceedings of the Annual Meeting of the Cognitive Science Society, vol. 46, 2024
2024
-
[30]
Analyzing the sensibility of visual language models using an evolving image generation system: Focusing on color impressions and sound symbolism,
R. Shinto and H. Iizuka, “Analyzing the sensibility of visual language models using an evolving image generation system: Focusing on color impressions and sound symbolism,” inArtificial Life Conference Pro- ceedings, vol. 36, 2024
2024
-
[31]
With ears to see and eyes to hear: Sound symbolism experiments with multimodal large language models,
T. Loakman, Y . Li, and C. Lin, “With ears to see and eyes to hear: Sound symbolism experiments with multimodal large language models,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024, pp. 2849–2867
2024
-
[32]
Sound to meaning mappings in the bouba-kiki effect,
K. McCormick, J. Y . Kim, S. List, and L. C. Nygaard, “Sound to meaning mappings in the bouba-kiki effect,” inProceedings of the 37th Annual Conference of the Cognitive Science Society, 2015, pp. 1565– 1570
2015
-
[33]
Correlation calculated from faulty data,
C. Spearman, “Correlation calculated from faulty data,”British journal of psychology, vol. 3, no. 3, p. 271, 1910
1910
-
[34]
Some experimental results in the correlation of mental abilities 1,
W. Brown, “Some experimental results in the correlation of mental abilities 1,”British Journal of Psychology, 1904-1920, vol. 3, no. 3, pp. 296–322, 1910
1904
-
[35]
The detection of disease clustering and a generalized regression approach,
N. Mantel, “The detection of disease clustering and a generalized regression approach,”Cancer research, vol. 27, no. 2 Part 1, pp. 209– 220, 1967
1967
-
[36]
Gemma 4 model card,
Google DeepMind, “Gemma 4 model card,” https://ai.google.dev/gemm a/docs/core/model card 4, 2026. APPENDIXA ETHICALCONSIDERATIONS Human subjects.This study does not collect new data from human participants. All human perceptual data used here were collected by prior research teams under their own institutional review procedures. No IRB approval was requir...
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.