REVIEW 3 major objections 2 minor 25 references
semantic-features: A User-Friendly Tool for Studying Contextual Word Embeddings in Interpretable Semantic Spaces
T0 review · 3 major / 2 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Masked language models assign a dative recipient higher animacy in the double-object construction and higher place-hood in the prepositional construction.
desk verdict A genuinely useful tool library with a plausible but unvalidated case study; the projector's generalization on individual contexts is the key missing piece. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a trained multi-layer perceptron that maps a contextual word embedding to a vector of Binder feature values. The network is trained on context-averaged embeddings from the British National Corpus against human-annotated feature norms; once trained, it can be applied to the embedding of any target word in a new sentence. The Binder space is the key because its features have concrete definitions that cleanly separate person-hood (Biomotion, Body, Human, Face, Speech) from place-hood (Landmark, Scene), allowing the authors to quantify how the model's representation shifts across constructions.
What would settle it
Collect human judgments of animacy and place-hood for the 450 experimental sentences, then compare them to the projected feature scores; if the projections do not align with human intuitions, or if the same names in neutral contexts also show the same projection pattern, the claimed construction-specific sensitivity would be called into question.
Extended reading notes
Core claim
The central empirical discovery is that the contextual representation of an ambiguous proper noun changes with the dative construction: projection into the Binder feature space—a set of human-annotated semantic properties such as Human, Landmark, and Scene—shows a rise in person/animacy features (Human, Body, Face, Biomotion, Speech) for the recipient in the double-object construction, and a rise in place features (Landmark, Scene) in the prepositional construction. This pattern holds for BERT, RoBERTa, and ALBERT across most layers, with the strongest effects in intermediate layers (6–9). The paper argues this indicates that contextually sensitive distributional embeddings capture subtle construction-level semantic changes.
Load-bearing premise
The trained projectors, fit on context-averaged embeddings from a general corpus, are assumed to generalize faithfully to the individual, unseen contexts of the experimental sentences, yet the paper does not evaluate projection accuracy on those specific stimuli.
Editorial extensions
If this is right
- The library can be used to test whether other grammatical constructions that impose semantic constraints on arguments, such as locative alternations or passive voice, are reflected in contextual embeddings.
- The finding that middle layers (6–9) carry the strongest semantic sensitivity suggests that layer-wise probing with interpretable spaces can localize where constructional semantics are encoded.
- The open-source tool and interactive demo lower the barrier for linguists and cognitive scientists to run hypothesis-driven semantic analyses on transformer LMs without writing custom code.
- Because the method works with any masked LM, it provides a standardized route to compare the semantic competences of future models across languages and architectures.
Reading between the lines
- The observed double-object versus prepositional-object difference might partly be an artifact of the projection probe rather than a genuine LM semantic representation; a probe trained on averaged contexts could systematically bias ambiguous names toward one reading. Testing the same sentences with a probe trained on randomized context embeddings would help isolate the LM's contribution.
- The method's success on dative recipients suggests a general strategy for diagnosing 'constructional meaning' in LMs: choose a construction whose alternation has a clear semantic contrast, build a balanced minimal pair dataset, and read off the relevant feature values. This could be extended to other argument-structure alternations, such as the spray/load alternation or the causative/inchoative al
- If the effect is robust, it implies that LMs do not treat words as having a single fixed sense but dynamically adjust their semantic features to fit the argument structure—a property that could be exploited for controlled paraphrase generation or for detecting when models rely on surface form rather than meaning.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces semantic-features, a library for projecting contextual word embeddings from masked language models into interpretable semantic feature spaces such as the Binder norms, together with an interactive Gradio demo. The central case study tests whether the dative alternation affects the semantic construal of the recipient argument: for sentence pairs like 'I sent London the letter' (DO) versus 'I sent the letter to London' (PO), the authors hypothesize that the recipient is more animate in the DO and more place-like in the PO. Using projectors trained on BNC context-averaged embeddings, they report that BERT, RoBERTa, and ALBERT show the expected direction of change in 33 of 36 model-layer combinations (Figure 2).
Significance. If the empirical result holds, the paper provides both a useful open-source tool for interpretable analysis of contextual embeddings and evidence that masked language models encode construction-specific semantic constraints, aligning with theoretical accounts of the dative alternation. The paper's explicit strengths are the public code repository, the reproducible training pipeline for the projection models, and the accessible interactive demo. However, the validity of the central empirical claim depends critically on whether the trained projectors generalize to individual unseen contexts, which is not demonstrated; the current evidence is also weakened by the small effective stimulus sample and the absence of statistical uncertainty quantification.
major comments (3)
- [Section 3, stimulus construction and Figure 2] The central empirical claim depends on the trained MLP projectors faithfully mapping individual contextual embeddings from the experimental sentences into Binder feature space, but the projectors are trained only on context-averaged embeddings per word from BNC (Section 2), and no validation is reported on held-out individual contexts. For a word like 'London', the training target is a single static Binder norm vector, so the projector may learn a lexical mapping that is biased toward the dominant reading rather than sensitive to construction-level context. A concrete test would be to evaluate projected feature values for the 450 experimental sentences against human feature-annotation judgments, or at least to report projector accuracy on held-out individual contexts from a corpus, and to show that the DO-PO difference is not reproduced by a projector trained on shuffled or context-free inputs. Without such validation, the observed differences in Figure 2 could be artifacts of the probe rather than genuine LM semantic sensitivity.
- [Section 3, Results] The 450 sentence pairs are generated from only 15 recipient names, 6 verbs, and 5 agents, so the effective independent sample is much smaller than 450. The claim that the pattern holds in '33 of 36' model-layer combinations therefore overstates the strength of the evidence: the 15 recipient names are the critical random factor, and each name contributes 30 sentence pairs. The paper should report by-item variance, include a permutation or mixed-effects analysis that treats recipient name as a random effect, and provide confidence intervals or significance tests for the average changes shown in Figure 2. Without this, the observed consistency across layers could reflect idiosyncratic properties of a few names rather than a general phenomenon.
- [Section 2, Model training; Appendix B] The training procedure reports an 80-20 train-validation split and selection by validation loss, but no held-out test evaluation is described for the projection models. The fact that all 117 models used 2 layers with 50% dropout and early stopping (Appendix B) does not by itself establish that the learned projections are accurate. Since the case study interprets absolute differences in predicted feature values, the paper should report at least one quantitative accuracy measure (e.g., correlation or mean absolute error) on a held-out set of individual contexts, and ideally on the same classes of ambiguous person/place nouns used in the experiment.
minor comments (2)
- [Section 1 and Section 3] There are several typos: 'incontext' should be 'in context', 'a interpetable' should be 'an interpretable', and 'To what extend' in Section 3 should be 'To what extent'.
- [Section 3, Table 1 and Table 2] The hand-selected feature subsets for person-hood and place-hood are a central design choice, but the paper does not discuss how sensitive the results are to adding or removing individual features (e.g., excluding 'Biomotion' or including 'Vision'). A brief sensitivity analysis would strengthen confidence that the animacy and place-hood contrasts are not driven by a single feature.
Circularity Check
No significant circularity: the dative-case result is an external behavioral benchmark, and the projection models are fit to independent Binder norms from averaged BNC embeddings.
full rationale
The paper's derivation chain is not circular. The MLP projectors in Section 2 are trained to map context-averaged word embeddings extracted from the British National Corpus to static feature norms from Binder et al. (2016), an external source; the training target is not the DO/PO animacy contrast. The case study in Section 3 then applies these fixed projectors to individual contextual embeddings of recipients in newly generated dative sentences and compares projected feature values between the DO and PO frames. The reported difference is therefore computed from the LM's contextual embeddings through a projector learned on external norms, not defined by the experimental hypothesis. The feature sets for personhood and placehood are selected from Binder et al.'s published definitions, independently of the LM outputs. The self-citations to Chronis et al. (2023) and to minicons supply method provenance and a prior layer-level observation; neither reduces the dative result to the cited work, and the central claim would stand or fall on the Figure 2 measurements even if those citations were removed. Concerns about unvalidated projector generalization or stimulus non-independence are questions of experimental validity and statistical strength, not circularity of the derivation.
Assumptions & free parameters
free parameters (4)
- MLP projection parameters =
Not reported numerically; trained weights available via demo and models
- Optuna-selected hyperparameters =
Not reported per final model; search ranges in Table 3
- Hand-selected Person and Place feature subsets =
Person: Biomotion, Body, Human, Face, Speech; Place: Landmark, Scene
- Stimulus set composition =
15 ambiguous names, 6 dative verbs, 5 agents, 450 template pairs
assumptions (6)
- domain assumption Binder et al. (2016) feature norms are a valid interpretable semantic target space.
- domain assumption Averaging a word's contextual embeddings over BNC preserves enough information to train a projection that works for individual contexts.
- domain assumption The trained MLP projectors generalize to unseen contextual occurrences of words like 'London'.
- domain assumption Masked language model layerwise embeddings are appropriate substrates for contextual semantic interpretation.
- ad hoc to paper The hand-selected feature subsets cleanly separate personhood from placehood.
- domain assumption The 15 ChatGPT-generated names are genuinely ambiguous between person and place for the sentence templates.
Cite this review
Pith. "Pith review of semantic-features: A User-Friendly Tool for Studying Contextual Word Embeddings in Interpretable Semantic Spaces." pith.science (2026). https://pith.science/paper/VFDA7RVV
@misc{pith2026250606169,
author = {Pith},
title = {Pith review of: semantic-features: A User-Friendly Tool for Studying Contextual Word Embeddings in Interpretable Semantic Spaces},
year = {2026},
howpublished = {\url{https://pith.science/paper/VFDA7RVV}},
note = {Machine review of arXiv:2506.06169}
}
read the original abstract
We introduce semantic-features, an extensible, easy-to-use library based on Chronis et al. (2023) for studying contextualized word embeddings of LMs by projecting them into interpretable spaces. We apply this tool in an experiment where we measure the contextual effect of the choice of dative construction (prepositional or double object) on the semantic interpretation of utterances (Bresnan, 2007). Specifically, we test whether "London" in "I sent London the letter." is more likely to be interpreted as an animate referent (e.g., as the name of a person) than in "I sent the letter to London." To this end, we devise a dataset of 450 sentence pairs, one in each dative construction, with recipients being ambiguous with respect to person-hood vs. place-hood. By applying semantic-features, we show that the contextualized word embeddings of three masked language models show the expected sensitivities. This leaves us optimistic about the usefulness of our tool.
Figures
Reference graph
Works this paper leans on
-
[1]
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. 2019. http://arxiv.org/abs/1907.10902 Optuna: A next-generation hyperparameter optimization framework
arXiv 2019
-
[2]
John Beavers. 2011. An aspectual analysis of ditransitive verbs of caused possession in english. Journal of semantics, 28(1):1--54
work page 2011
-
[3]
Jeffrey R. Binder, Lisa L. Conant, Colin J. Humphries, Leonardo Fernandino, Stephen B. Simons, Mario Aguilar, and Rutvik H. Desai. 2016. https://doi.org/10.1080/02643294.2016.1147426 Toward a brain-based componential semantic representation . Cognitive Neuropsychology, 33(3):130--174
arXiv 2016
-
[4]
Joan Bresnan. 2007. Predicting the dative alternation. Cognitive foundations of interpretation/Royal Netherlands Academy of Science
work page 2007
-
[5]
Erin M. Buchanan, K. D. Valentine, and Nicholas P. Maxwell. 2019. https://doi.org/10.3758/s13428-019-01243-z English semantic feature production norms: An extended database of 4436 concepts . Behavior Research Methods, 51:1849--1863
-
[6]
Gabriella Chronis, Kyle Mahowald, and Katrin Erk. 2023. https://doi.org/10.18653/v1/2023.acl-long.14 A method for studying semantic construal in grammatical constructions with interpretable contextual embedding spaces . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 242--261, Toron...
-
[7]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long a...
-
[8]
Katrin Erk. 2009. https://aclanthology.org/W09-1109 Representing words as regions in vector space . In Proceedings of the Thirteenth Conference on Computational Natural Language Learning ( C o NLL -2009) , pages 57--65, Boulder, Colorado. Association for Computational Linguistics
work page 2009
Show all 25 references
-
[9]
Adele E Goldberg. 1995. Constructions: A Construction Grammar Approach to Argument Structure. University of Chicago Press
1995
-
[10]
Robert Hawkins, Takateru Yamakoshi, Thomas Griffiths, and Adele Goldberg. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.376 Investigating representations of verb bias in neural language models . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language...
2020 doi
-
[11]
Malka Rappaport Hovav and Beth Levin. 2008. The english dative alternation: The case for verb sensitivity. Journal of linguistics, 44(1):129--167
2008
-
[12]
Jaap Jumelet, Willem Zuidema, and Arabella Sinclair. 2024. https://doi.org/10.18653/v1/2024.findings-acl.877 Do language models exhibit human-like structural priming effects? In Findings of the Association for Computational Linguistics: ACL 2024, pages 14727--14742, Bangkok, T...
2024 doi
-
[13]
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations . In International Conference on Learning Representations
2020
-
[14]
Beth Levin. 1993. English verb classes and alternations: A preliminary investigation. University of Chicago press
1993
-
[15]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. http://arxiv.org/abs/1907.11692 RoBERTa : A Robustly Optimized BERT Pretraining Approach . arXiv:1907.11692 [cs]. ArXiv: 1907.11692
2019 arXiv
-
[16]
Zoey Liu and Stefanie Wulff. 2023. https://aclanthology.org/2023.depling-1.1 The development of dependency length minimization in early child language: A case study of the dative alternation . In Proceedings of the Seventh International Conference on Dependency Linguistics (De...
2023
-
[17]
Ken McRae, George S Cree, Mark S Seidenberg, and Chris McNorgan. 2005. Semantic feature production norms for a large set of living and nonliving things. Behavior research methods, 37(4):547--559
2005
-
[18]
Tomas Mikolov , Kai Chen , Greg Corrado , and Jeffrey Dean . 2013. http://arxiv.org/abs/1301.3781 Efficient Estimation of Word Representations in Vector Space . arXiv e-prints, page arXiv:1301.3781
2013 arXiv
-
[19]
Kanishka Misra. 2022. https://arxiv.org/abs/2203.13112 minicons: Enabling flexible behavioral and representational analyses of transformer language models . arXiv:2203.13112
2022 arXiv
-
[20]
Kanishka Misra and Najoung Kim. 2024. Generating novel experimental hypotheses from language models: A case study on cross-dative generalization. arXiv preprint arXiv:2408.05086
2024 arXiv
-
[21]
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. Glove: Global vectors for word representation. In Proceedings of EMNLP 2014 , pages 1532--1543
2014
-
[22]
Jackson Petty, Michael Wilson, and Robert Frank. 2022. Do language models learn position-role mappings? In Proceedings of the 46th annual Boston University Conference on Language Development
2022
-
[23]
Qing Yao, Kanishka Misra, Leonie Weissweiler, and Kyle Mahowald. 2025. Both direct and indirect evidence contribute to dative alternation preferences in language models. arXiv preprint arXiv:2503.20850
2025 arXiv
-
[24]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before...
-
[25]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.