Pith. sign in

REVIEW 4 major objections 6 minor 243 references

Relation Geometry in Semantic Space of Language Models

T0 review · 4 major / 6 minor · reviewed 2026-07-30 · grok-4.5

Pith's one-line read Semantic relations do not leave equally clear footprints in language-model vector space; asymmetric ones are clearer than symmetric ones.

desk verdict Solid multi-relation geometry audit with unusually clean anti-leakage design; the asymmetric-vs-symmetric pattern is real, but absolute scores are modest and the distributional-hypothesis punchline outruns the linear-probe evidence. read the letter →

arxiv 2607.26762 v1 pith:YKLCJEPB submitted 2026-07-29 cs.CL

classification cs.CL
keywords semanticrelationsrelationgeometrylanguagemodelsprobingdistributionalsemanticshypernymysynonymylexicalvscontextualinformation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Language models turn words into vectors, but it is unclear how much of the structure of semantic relations—hypernymy, synonymy, antonymy, and the rest—is actually present as geometry in those spaces. This paper tests that question directly: whether words related to a target by the same relation cluster in a shared region, whether those regions encode classic properties such as asymmetry and transitivity, and whether the signal comes more from word form or from context. Across causal, masked, and diffusion models, asymmetric relations show relatively clearer separable regions than symmetric ones, while directionality and long-distance transitivity are only moderately recovered. The pattern is uneven enough to suggest that not every semantic relation is equally learnable from distributional information alone. A sympathetic reader cares because this is a concrete check on a foundational claim of distributional semantics, not just another leaderboard number.

What carries the argument

A bilinear multi-relation probe that scores whether a target–relatum pair falls into a relation-specific linear “relata region,” evaluated with controlled predictability, directionality, and transitivity metrics, plus ablations that strip lexical form or correct context.

What would settle it

Under the same lemma-held-out, leakage-blocked setup, synonym or antonym pairs becoming as linearly separable and directionally consistent as hypernym or hyponym pairs, or long-distance hypernym chains remaining highly predictable when the probe is trained only on direct pairs.

Watch

Extended reading notes

Core claim

Relation geometry in language-model semantic space is not uniform: relata of asymmetric relations occupy relatively distinct linear regions, while symmetric relations do not, and properties such as directionality and especially long-distance transitivity are only moderately encoded—evidence that semantic relations are not equally recoverable from distributional learning alone.

Load-bearing premise

That how well this linear probe separates held-out word pairs is a faithful readout of whether relation-specific regions and relation properties actually exist in the model’s space.

Editorial extensions

If this is right

  • Asymmetric relatedness (hypernymy, meronymy and their reverses) should be easier to read out of LM embeddings than synonymy-style similarity.
  • Probes and applications that assume uniform relational geometry across WordNet-style relations will systematically overestimate what distributional spaces encode for synonyms and antonyms.
  • Causal models will lean more on surface form for relational geometry, while masked and diffusion models will lean more on context—except for indirect hypernymy, where lexical form matters more across model types.
  • Long-distance hierarchical inference cannot be treated as a free consequence of local hypernym geometry learned from co-occurrence alone.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If linear relata regions are weak for similarity but stronger for relatedness, many “analogy” or vector-offset recipes may be succeeding on relatedness structure rather than true synonymy.
  • The gap between short- and long-distance transitivity suggests hierarchical resources still need explicit structure beyond what next-token or masked training induces.
  • A natural next test is whether non-linear probes close the synonymy gap or merely confirm that the missing structure is absent, not just hard to read linearly.
  • Training objectives that mix shared-neighbour and co-occurrence signals may explain why LMs beat static embeddings on asymmetric relations but not on antonymy.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper asks whether six WordNet semantic relations are encoded as linearly separable “relata regions” in LM token spaces, and whether those regions reflect symmetry/asymmetry and hypernymy transitivity. It trains bilinear multi-relation probes (Eqs. 2–4) on final-layer representations from ModernBERT, LLaDA-8B, LLaMA-3.1-8B, and a fastText baseline, using SemCor/WordNet triplets with lemma-held-out folds, removal of property-exhibiting pairs from training, an unrelated control class, random-representation controls, and multi-trial CV with Bonferroni-corrected tests. It further ablates lexical vs. contextual information via masking/previous-token states and degenerate/egalitarian attention. Controlled scores show clearer (still modest) separability for asymmetric than symmetric relations, only moderate directionality and short-distance transitivity, and model-dependent reliance on lexical vs. contextual cues. The authors read this as evidence that relation geometry is uneven and that semantic relations are not uniformly learnable from distributional information alone.

Significance. The work usefully broadens relation probing beyond hypernymy, introduces property-aware metrics (directionality; leakage-blocked transitivity), a geometrically interpretable bilinear probe, and a quantitative lexical/contextual ablation (ΔLEX/ΔCTX). Methodological care—lemma anti-leakage splits, property-pair removal, random probes, multi-trial CV, and transparent reporting of low absolute controlled scores—is a genuine strength and raises the bar relative to much of the prior probing literature. If the relative asymmetric/symmetric pattern and the information-source differences hold under tighter controls, the paper supplies concrete empirical constraints for distributional-semantics theory and for how different LM objectives encode relational structure. The contribution is primarily empirical and methodological rather than a decisive theoretical refutation.

major comments (4)
  1. [Abstract; §5.1–5.3; §5.5–6; Tables 4–6; Fig. 2] The central theoretical suggestion (abstract; §5.5–6)—that uneven relation geometry shows semantic relations are not uniformly learnable from distribution alone—overreaches the absolute evidence. Controlled F_r never exceeds ~0.27 (asymmetric) / ~0.16 (symmetric) (Table 4); D_r ≤ 0.18 (Table 5); long-distance T falls to chance or below (Fig. 2, Table 6). Relative gaps are real under the authors’ controls, but without a positive control showing that the same bilinear probe recovers high controlled scores when linear relata-region geometry is independently known to be present, modest absolute performance cannot securely license a counter-claim against distributional semantics. Temper the claim to what the relative pattern supports, or add such a control.
  2. [Table 3; §5.1–5.4; §5.5; §6–7] Model comparison confounds scale, data, and objective (Table 3: ModernBERT 395M vs LLaMA/LLaDA 8B; unmatched corpora/context windows). LLaMA’s advantage on asymmetric F/D/T is repeatedly attributed in part to the causal objective and bidirectional vs unidirectional context (§5.5), yet Limitations §7 correctly notes these factors are entangled. As written, the abstract and discussion still invite a family-level conclusion (CLM vs MLM/DLM; lexical vs contextual importance). Either match models more carefully, add ablations that isolate objective/scale, or systematically downgrade causal language about “causal vs masked/diffusion” throughout results and conclusion.
  3. [§3.1 Eqs. 2–5; §5.5–6; §7] The operational definition of “relation geometry” is linear separability under the bilinear probe (Eqs. 2–5, §3.1). Limitations §7 acknowledges that non-linear geometry is untested and that probing ≠ use. Because the negative findings on synonymy/antonymy and weak long-distance transitivity are load-bearing for the “not equally well-represented / not uniformly learnable” claim, the paper should either (i) include at least one non-linear probe baseline as a sensitivity check, or (ii) consistently frame all conclusions as about linear relata regions in final-layer space rather than about relation geometry or distributional learnability in general. The current framing oscillates between these.
  4. [§3.2.3; §4.4; Fig. 2; §5.3; §5.5] Transitivity evaluation trains only on direct pairs and tests indirect pairs (good anti-leakage design), but performance collapse with distance (Fig. 2) is also consistent with sense/representation drift and decreasing semantic overlap, not only with absence of transitive geometry. The discussion (§5.5) notes contextual dissimilarity for long-distance pairs; that alternative should be quantified (e.g., baseline similarity or probe confusion as a function of path length) so that “transitivity not encoded” is distinguished from “indirect pairs are simply harder under the same linear readout.”
minor comments (6)
  1. [Front matter] Placeholder metadata remains in the front matter (“Action editor: {action editor name}”; “Submission received: DD Month YYYY”). Clean before any revision cycle.
  2. [Headers] Running headers still say “Author’s Surnames Here / Running Article Title Here”.
  3. [§4.2; Table 1; Table 2] Table 1 token counts and the 6.8M indirect-triplet figure in the intro should be cross-checked for consistency with the filtering narrative (intra-sentential removal, distance cutoffs).
  4. [Figure 1; §3.1] Figure 1 is schematic only; a 2D illustration of real probe hyperplanes or a qualitative PCA/UMAP of one target’s relata would help readers judge what “region” means empirically.
  5. [§4.5; Appendix C] Appendix C layer analysis is valuable but under-discussed in the main text; a short pointer in §5 on whether upper-layer dominance holds for all metrics would strengthen the final-layer choice.
  6. [§3.2.1; §3.4] Minor wording: “summay relation predictabilities” (§3.2.1); “egalitarian decontextualisation” notation could be introduced once with a short pseudocode block for reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

Empirical probing study with external gold labels and held-out evaluation; no derivation reduces to its inputs by construction.

full rationale

The paper operationalizes ‘relation geometry’ via bilinear probe performance (Eqs. 2–9) on LM token pairs, with WordNet/SemCor providing external gold relations, lemma-held-out cross-validation, property-leakage blocking, and random-representation controls. Probe F/D/T scores and lexical/contextual ablations are measurements against that external structure, not quantities fitted from the same data and re-labeled as predictions. Bilinear scoring is imported from knowledge-graph embedding literature (Bordes et al., Nickel et al., etc.), not from a self-citation uniqueness claim. Self-citations (e.g. Cao et al. 2025 on antonymy) are peripheral. There is no self-definitional loop, no fitted-input-as-prediction step, and no load-bearing self-citation chain. Absolute scores are modest and models unmatched—these are validity/scope issues, not circularity. Derivation chain is self-contained empirical evaluation; score 0.

Assumptions & free parameters 3 free parameters · 5 assumptions · 2 invented entities

Load-bearing content is methodological and empirical rather than axiomatic physics-style theory. The claim rests on standard distributional-semantics background, WordNet as gold taxonomy, the assumption that bilinear separability indexes “relata regions,” and several experimental design choices (linear probes, final layer, noun-only SemCor, leakage filters). No new physical entities; free parameters are ordinary ML hyperparameters and design thresholds, not constants fitted to prove a law.

free parameters (3)
  • Attention temperature τ for egalitarian decontextualisation = 100
    Set by hand to 100 after informal search so attention is effectively uniform; enters all ΔCTX comparisons.
  • Probe training hyperparameters (Adam, 5 epochs, batch 1024, 4-fold × 10 trials) = 5 epochs, bs=1024, 4×10 CV
    Chosen experimental settings that determine absolute probe fit; controlled scores mitigate but do not remove dependence.
  • UNR sample size and max hypernymy distance cutoff = 500000 UNR; distance ≤9 kept
    500k unrelated triplets sampled; distances >10 and non-unique distances discarded (~5% indirect loss), shaping class balance and T metrics.
assumptions (5)
  • domain assumption Harris-style distributional semantics: regular differences in environments correspond to semantic relations that could appear as geometry in embedding space.
    Frames the entire research question in §1; the paper tests rather than proves it.
  • domain assumption WordNet synset relations on SemCor-disambiguated nouns are an adequate gold standard for the six relations and for hypernym path distance.
    All triplets and transitivity distances are defined from WordNet (§4); known WordNet DAG ambiguities are only partly filtered.
  • ad hoc to paper A softmax over bilinear scores e_w^T W_r e_v measures linear separability of relation-specific relata regions in the LM’s token space.
    Core operationalisation in §3.1; alternative combination ops or nonlinear probes could yield different geometry conclusions (acknowledged in §7).
  • ad hoc to paper Random vectors drawn Uniform[0,1] yield an appropriate chance baseline so that controlled performance isolates geometry rather than label skew.
    §3.3 control procedure; baseline distribution choice is conventional but not uniquely justified.
  • ad hoc to paper Removing the surface token (mask or previous-token state) vs forcing degenerate/egalitarian attention cleanly isolates lexical vs contextual contributions.
    §3.4 ablation design; entanglement of form and context in real LMs means isolation is approximate.
invented entities (2)
  • Relata regions
    purpose: Name the hypothesized linear regions occupied by words in a fixed relation to a target.
    Central geometric construct illustrated in Fig. 1 and tested via probes; descriptive rather than a new ontological object in nature.
  • Controlled performance / ΔLEX / ΔCTX loss metrics
    purpose: Quantify geometry beyond raw probe accuracy and importance of lexical vs contextual information.
    Paper-defined evaluation quantities; useful instrumentation, not entities with external existence claims.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Relation Geometry in Semantic Space of Language Models." pith.science (2026). https://pith.science/paper/YKLCJEPB

@misc{pith2026260726762,
  author       = {Pith},
  title        = {Pith review of: Relation Geometry in Semantic Space of Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YKLCJEPB}},
  note         = {Machine review of arXiv:2607.26762}
}
read the original abstract

When it comes to generating vector representations of words, current language models are achieving high-quality results. However, what is not known is the extent to which knowledge about semantic relations is represented in the geometry of the semantic spaces created in this way. In order to answer this question, we study the relation geometry of such semantic spaces from three perspectives. We first examine whether words standing in a particular relation to a target word~(called relata) occupy the same region in semantic space, and whether the regions corresponding to different relations are distinct from each other. We then verify to what extent semantic spaces reflect certain well-known properties of relations, such as symmetry, asymmetry, and transitivity. Finally, we consider which information about the target words and relata is more important for relation geometry: their surface forms, or their contexts. We conduct experiments on six semantic relations using causal, masked, and diffusion language models. The results show that relata in asymmetric relations relatively clearly occupy a distinct region in semantic space. Asymmetric relations' properties are only moderately well encoded in the semantic space, yet better than those of symmetric ones. Furthermore, when considering the question which information source has the strongest impact on results amongst the models we evaluated, we find that lexical information tends to be more important for the causal language model, whereas contextual information is more important for the masked and diffusion language models. Our results empirically show that relation geometry is not equally well-represented for all relations in semantic space, suggesting that there is a difference in how well semantic relations might be learned from distributional information alone.

Figures

Figures reproduced from arXiv: 2607.26762 by the authors.

Figure 1
Figure 1. If this is true, a simple prediction task should be able to reveal this correspon [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

243 extracted references · 84 canonical work pages

  1. [1]

    No clues good clues: out of context Lexical Relation Classification

    Pitarch, Lucia and Bernad, Jordi and Dranca, Lacramioara and Bobed Lisbona, Carlos and Gracia, Jorge. No clues good clues: out of context Lexical Relation Classification. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.308

  2. [3]

    Inclusive yet Selective: Supervised Distributional Hypernymy Detection

    Roller, Stephen and Erk, Katrin and Boleda, Gemma. Inclusive yet Selective: Supervised Distributional Hypernymy Detection. Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers. 2014

  3. [4]

    Distributional Inclusion Hypothesis and Quantifications: Probing for Hypernymy in Functional Distributional Semantics

    Lo, Chun Hei and Lam, Wai and Cheng, Hong and Emerson, Guy. Distributional Inclusion Hypothesis and Quantifications: Probing for Hypernymy in Functional Distributional Semantics. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. doi:10.18653/v1/2024.acl-long.784

  4. [5]

    Introducing Orthogonal Constraint in Structural Probes

    Limisiewicz, Tomasz and Mare c ek, David. Introducing Orthogonal Constraint in Structural Probes. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. doi:10.18653/v1/2021.acl-long.36

  5. [6]

    P ro SA : Assessing and Understanding the Prompt Sensitivity of LLM s

    Zhuo, Jingming and Zhang, Songyang and Fang, Xinyu and Duan, Haodong and Lin, Dahua and Chen, Kai. P ro SA : Assessing and Understanding the Prompt Sensitivity of LLM s. Findings of the Association for Computational Linguistics: EMNLP 2024. 2024. doi:10.18653/v1/2024.findings-emnlp.108

  6. [7]

    What Don ' t RNN Language Models Learn About Filler-Gap Dependencies?

    Chaves, Rui. What Don ' t RNN Language Models Learn About Filler-Gap Dependencies?. Proceedings of the Society for Computation in Linguistics 2020. 2020

  7. [8]

    Computational Linguistics , author =

    Language. Computational Linguistics , author =. 2024 , keywords =. doi:10.1162/coli_a_00492 , number =

  8. [9]

    Frontiers of Computer Science , author =

    A survey on large language model based autonomous agents , volume =. Frontiers of Computer Science , author =. 2024 , pages =. doi:10.1007/s11704-024-40231-1 , number =

Show all 243 references
  1. [10]

    CoRR , volume =

    Tianyi Li and Mingda Chen and Bowei Guo and Zhiqiang Shen , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2508.10875 , eprinttype =. 2508.10875 , timestamp =

  2. [11]

    Large Language Diffusion Models , journal =

    Shen Nie and Fengqi Zhu and Zebin You and Xiaolu Zhang and Jingyang Ou and Jun Hu and Jun Zhou and Yankai Lin and Ji. Large Language Diffusion Models , journal =. 2025 , url =. doi:10.48550/ARXIV.2502.09992 , eprinttype =. 2502.09992 , timestamp =

  3. [12]

    Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

    Warner, Benjamin and Chaffin, Antoine and Clavi. Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volu...

  4. [13]

    Word , author =

    Distributional structure , volume =. Word , author =. 1954 , keywords =. doi:10.1080/00437956.1954.11659520 , abstract =

  5. [14]

    Firth, J. R. , year =. A synopsis of linguistic theory, 1930-1955 , journal =

  6. [15]

    The Vector Grounding Problem , journal =

    Dimitri Coelho Mollo and Rapha. The Vector Grounding Problem , journal =. 2023 , url =. doi:10.48550/ARXIV.2304.01481 , eprinttype =. 2304.01481 , timestamp =

  7. [16]

    Physica D: Nonlinear Phenomena , volume =

    Stevan Harnad , title =. Physica D: Nonlinear Phenomena , volume =. 1990 , url =

  8. [17]

    The Italian Journal of Linguistics , year=

    The Distributional Hypothesis , author=. The Italian Journal of Linguistics , year=

  9. [18]

    Weinberger and Yoav Artzi , title =

    Tianyi Zhang and Varsha Kishore and Felix Wu and Kilian Q. Weinberger and Yoav Artzi , title =. 8th International Conference on Learning Representations,. 2020 , url =

  10. [19]

    A Fine-Grained Analysis of BERTS core

    Hanna, Michael and Bojar, Ond r ej. A Fine-Grained Analysis of BERTS core. Proceedings of the Sixth Conference on Machine Translation. 2021

  11. [20]

    Varshney and Caiming Xiong and Richard Socher , title =

    Nitish Shirish Keskar and Bryan McCann and Lav R. Varshney and Caiming Xiong and Richard Socher , title =. CoRR , volume =. 2019 , url =. 1909.05858 , timestamp =

  12. [21]

    Neural Word Embedding as Implicit Matrix Factorization , url =

    Levy, Omer and Goldberg, Yoav , booktitle =. Neural Word Embedding as Implicit Matrix Factorization , url =

  13. [22]

    and Furnas, George W

    Deerwester, Scott and Dumais, Susan T. and Furnas, George W. and Landauer, Thomas K. and Harshman, Richard , title =. Journal of the American Society for Information Science , volume =. doi:https://doi.org/10.1002/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO;2-9 , url =. h...

  14. [23]

    Attention is All you Need , url =

    Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N and Kaiser, ukasz and Polosukhin, Illia , booktitle =. Attention is All you Need , url =

  15. [24]

    Learning Phrase Representations using RNN Encoder -- Decoder for Statistical Machine Translation

    Cho, Kyunghyun and van Merri. Learning Phrase Representations using RNN Encoder -- Decoder for Statistical Machine Translation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing ( EMNLP ). 2014. doi:10.3115/v1/D14-1179

  16. [25]

    A Neural Probabilistic Language Model , url =

    Bengio, Yoshua and Ducharme, R\'. A Neural Probabilistic Language Model , url =. Advances in Neural Information Processing Systems , editor =

  17. [26]

    Discover Computing , author =

    A unified and scalable machine learning framework for feature fusion in object classification using weighted. Discover Computing , author =. 2025 , pages =. doi:10.1007/s10791-025-09622-1 , number =

  18. [27]

    2024 , pages =

    Neural Processing Letters , author =. 2024 , pages =. doi:10.1007/s11063-024-11464-9 , number =

  19. [28]

    Does BERT Know that the IS -A Relation Is Transitive?

    Lin, Ruixi and Ng, Hwee Tou. Does BERT Know that the IS -A Relation Is Transitive?. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2022. doi:10.18653/v1/2022.acl-short.11

  20. [29]

    Exploring the Representation of Word Meanings in Context: A Case Study on Homonymy and Synonymy

    Garcia, Marcos. Exploring the Representation of Word Meanings in Context: A Case Study on Homonymy and Synonymy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (...

  21. [30]

    Bridging Perception, Memory, and Inference through Semantic Relations

    Bj. Bridging Perception, Memory, and Inference through Semantic Relations. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. doi:10.18653/v1/2021.emnlp-main.719

  22. [31]

    Inspecting the concept knowledge graph encoded by modern language models

    Aspillaga, Carlos and Mendoza, Marcelo and Soto, Alvaro. Inspecting the concept knowledge graph encoded by modern language models. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. 2021. doi:10.18653/v1/2021.findings-acl.263

  23. [32]

    R eliable E val: A Recipe for Stochastic LLM Evaluation via Method of Moments

    Lior, Gili and Habba, Eliya and Levy, Shahar and Caciularu, Avi and Stanovsky, Gabriel. R eliable E val: A Recipe for Stochastic LLM Evaluation via Method of Moments. Findings of the Association for Computational Linguistics: EMNLP 2025. 2025. doi:10.18653/v1/2025.findings-emnlp.594

  24. [33]

    How Do LLM s Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training

    Ou, Yixin and Yao, Yunzhi and Zhang, Ningyu and Jin, Hui and Sun, Jiacheng and Deng, Shumin and Li, Zhenguo and Chen, Huajun. How Do LLM s Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training. Findings of the Association for Computational Linguisti...

  25. [34]

    CoRR , volume =

    Mattia Proietti and Alessandro Lenci , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2504.02395 , eprinttype =. 2504.02395 , timestamp =

  26. [35]

    On the Distinctive Co-occurrence Characteristics of Antonymy

    Cao, Zhihan and Yamada, Hiroaki and Tokunaga, Takenobu. On the Distinctive Co-occurrence Characteristics of Antonymy. Proceedings of the 14th Joint Conference on Lexical and Computational Semantics (*SEM 2025). 2025. doi:10.18653/v1/2025.starsem-1.10

  27. [36]

    Du, Mengnan and He, Fengxiang and Zou, Na and Tao, Dacheng and Hu, Xia , title =. Commun. ACM , month = dec, pages =. 2023 , issue_date =. doi:10.1145/3596490 , abstract =

  28. [37]

    Language Resources and Evaluation , year =

    Zhihan Cao and Hiroaki Yamada and Simone Teufel and Takenobu Tokunaga , title =. Language Resources and Evaluation , year =

  29. [38]

    and Wiersma, William and Jurs, Stephen G

    Hinkle, Dennis E. and Wiersma, William and Jurs, Stephen G. Applied statistics for the behavioral sciences. 2003

  30. [39]

    The Corpus of Contemporary American English (COCA)

    Davies, Mark. The Corpus of Contemporary American English (COCA). 2008

  31. [40]

    Global WordNet Conference 2025 , year =

    Zhihan Cao and Hiroaki Yamada and Simone Teufel and Takenobu Tokunaga , title =. Global WordNet Conference 2025 , year =

  32. [41]

    The LAMBADA dataset: Word prediction requiring a broad discourse context

    Paperno, Denis and Kruszewski, Germ \'a n and Lazaridou, Angeliki and Pham, Ngoc Quan and Bernardi, Raffaella and Pezzelle, Sandro and Baroni, Marco and Boleda, Gemma and Fern \'a ndez, Raquel. The LAMBADA dataset: Word prediction requiring a broad discourse context. Proceedin...

  33. [42]

    Brown and Benjamin Chess and Rewon Child and Scott Gray and Alec Radford and Jeffrey Wu and Dario Amodei , title =

    Jared Kaplan and Sam McCandlish and Tom Henighan and Tom B. Brown and Benjamin Chess and Rewon Child and Scott Gray and Alec Radford and Jeffrey Wu and Dario Amodei , title =. CoRR , volume =. 2020 , url =. 2001.08361 , timestamp =

  34. [43]

    and Chaffin, Roger and Herrmann, Douglas

    Winston, Morton E. and Chaffin, Roger and Herrmann, Douglas. A Taxonomy of Part-Whole Relations. Cognitive Science. 1987. doi:https://doi.org/10.1207/s15516709cog1104\_2

  35. [44]

    and Bousfield, Weston A

    Cohen, Burton H. and Bousfield, Weston A. and Whitmarsh,Geneva. Cultural Norms For Verba Items in 43 Categories. Studies on the Mediation of Verbal Behavior: Technical Report. 1957

  36. [45]

    Where Partonomies and Taxonomies Meet

    Tversky, Barbara. Where Partonomies and Taxonomies Meet. Meanings and Prototypes (RLE Linguistics B: Grammar): Studies in Linguistic Categorization (1st ed.). 2014

  37. [46]

    Possessives in English: An Exploration in Cognitive Grammar

    Taylor, John R. Possessives in English: An Exploration in Cognitive Grammar. 1996. doi:10.1093/oso/9780198235866.001.0001

  38. [47]

    Alan Cruse

    D. Alan Cruse. Lexical Semantics. 1986

  39. [48]

    Alan Cruse

    D. Alan Cruse. Prototype theory and lexical relations. Rivista di linguistica. 1994

  40. [49]

    Noms collectifs et méronymie

    Lecolle, Michelle. Noms collectifs et méronymie. Cahiers de grammaire. 1998

  41. [50]

    Wieso ist ein Kollektivum ein Kollektivum? Zentrum und Peripherieeiner Kategorie am Beispiel des Spanischen

    Mihatsch, Wiltrud. Wieso ist ein Kollektivum ein Kollektivum? Zentrum und Peripherieeiner Kategorie am Beispiel des Spanischen. Philologie im Netz. 2000

  42. [51]

    Marry L. McHugh. Interrater reliability: the kappa statistic. Biochemia Medica. 2012. doi:10.11613/BM.2012.031

  43. [52]

    Measuring the Reliability of Qualitative Text Analysis Data

    Klaus krippendorff. Measuring the Reliability of Qualitative Text Analysis Data. Quality & Quantity. 2004. doi:10.1007/s11135-004-8107-7

  44. [53]

    Proceedings of the 43rd Annual Meeting on Association for Computational Linguistics , pages =

    Geffet, Maayan and Dagan, Ido , title =. Proceedings of the 43rd Annual Meeting on Association for Computational Linguistics , pages =. 2005 , publisher =. doi:10.3115/1219840.1219854 , abstract =

  45. [54]

    Proceedings of the International Conference on Language Resources and Evaluation (LREC 2018) , year=

    Learning Word Vectors for 157 Languages , author=. Proceedings of the International Conference on Language Resources and Evaluation (LREC 2018) , year=

  46. [55]

    Nelson (Winthrop Nelson) and Twaddell, W

    Kučera, Henry and Francis, W. Nelson (Winthrop Nelson) and Twaddell, W. F. (William Freeman) and Marckworth, Mary Lois and Bell, Laura M. and Carroll, John Bissell. Computational analysis of present-day American English. 1967

  47. [56]

    Language Models are Few-Shot Learners

    Brown, Tom and Mann, Benjamin and Ryder, Nick and Subbiah, Melanie and Kaplan, Jared D and Dhariwal, Prafulla and Neelakantan, Arvind and Shyam, Pranav and Sastry, Girish and Askell, Amanda and Agarwal, Sandhini and Herbert-Voss, Ariel and Krueger, Gretchen and Henighan, Tom a...

  48. [57]

    Probing Classifiers: Promises, Shortcomings, and Advances

    Belinkov, Yonatan. Probing Classifiers: Promises, Shortcomings, and Advances. Computational Linguistics. 2022. doi:10.1162/coli_a_00422

  49. [58]

    The Analysis of Synonymy and Antonymy in Discourse Relations: An Interpretable Modeling Approach

    Asela Reig Alamillo and David Torres Moreno and Eliseo Morales González and Mauricio Toledo Acosta and Antoine Taroni and Jorge Hermosillo Valadez. The Analysis of Synonymy and Antonymy in Discourse Relations: An Interpretable Modeling Approach. Computational Linguistics. 2023...

  50. [59]

    A Semantic Approach to Recognizing Textual Entailment

    Marta Tatu and Dan Moldovan. A Semantic Approach to Recognizing Textual Entailment. Proceedings of Human Language Technology Conference and Conference on Empirical Methods in Natural Language Processing. 2005

  51. [60]

    Nitin Madnani and Bonnie J. Dorr. Generating Phrasal and Sentential Paraphrases: A Survey of Data-Driven Methods. Computational Linguistics. 2010. doi:10.1162/COLI_A_00002

  52. [61]

    Simplifying Lexical Simplification: Do We Need Simplified Corpora?

    Goran Glavaš and Sanja Štajner. Simplifying Lexical Simplification: Do We Need Simplified Corpora?. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Shor...

  53. [62]

    Miller and Christiane Fellbaum

    George A. Miller and Christiane Fellbaum. Semantic networks of english. Cognition. 1991. doi:10.1016/0010-0277(91)90036-4

  54. [63]

    Antonym order in English and Chinese coordinate structures , url =

    Wu, Shuqiong and Zhang, Jie , doi =. Antonym order in English and Chinese coordinate structures , url =. Review of Cognitive Linguistics , keywords =

  55. [64]

    Charles and George A

    Walter G. Charles and George A. Miller. Contexts of antonymous adjectives. Applied Psycholinguistics. 1989. doi:10.1017/S0142716400008675

  56. [65]

    Co-Occurrence and Antonymy

    Christiane Fellbaum. Co-Occurrence and Antonymy. International Journal of Lexicography. 1995. doi:10.1093/ijl/8.4.281

  57. [66]

    Justeson and Slava M

    John S. Justeson and Slava M. Katz. Co-occurrences of Antonymous Adjectives and Their Contexts. Computational Linguistics. 1991

  58. [67]

    Miller and Walter G

    George A. Miller and Walter G. Charles. Contextual correlates of semantic similarity. Language and Cognitive Processes. 1991. doi:10.1080/01690969108406936

  59. [68]

    Antonymy:

    Jones, Steven Jeffrey , year =. Antonymy:

  60. [69]

    Biometrical Journal , volume =

    Brunner, Edgar and Munzel, Ullrich , title =. Biometrical Journal , volume =. doi:https://doi.org/10.1002/(SICI)1521-4036(200001)42:1<17::AID-BIMJ17>3.0.CO;2-U , url =. https://onlinelibrary.wiley.com/doi/pdf/10.1002/ year =

  61. [70]

    How Do Large Language Models Acquire Factual Knowledge During Pretraining? , url =

    Chang, Hoyeon and Park, Jinho and Ye, Seonghyeon and Yang, Sohee and Seo, Youngkyung and Chang, Du-Seong and Seo, Minjoon , booktitle =. How Do Large Language Models Acquire Factual Knowledge During Pretraining? , url =

  62. [71]

    The Thirteenth International Conference on Learning Representations , year=

    Knowledge Localization: Mission Not Accomplished? Enter Query Localization! , author=. The Thirteenth International Conference on Learning Representations , year=

  63. [72]

    Dual Tensor Model for Detecting Asymmetric Lexico-Semantic Relations

    Goran Glavaš and Simone Paolo Ponzetto. Dual Tensor Model for Detecting Asymmetric Lexico-Semantic Relations. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 2017. doi:10.18653/v1/D17-1185

  64. [73]

    Antonym sequence in written discourse: a corpus-based study

    Nataša Kostić. Antonym sequence in written discourse: a corpus-based study. Language Sciences. 2015. doi:10.1016/j.langsci.2014.07.013

  65. [74]

    On Log-Likelihood-Ratios and the Significance of Rare Events

    Moore, Robert C. On Log-Likelihood-Ratios and the Significance of Rare Events. Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing. 2004

  66. [75]

    Corpora and collocations , volume =

    58. Corpora and collocations , volume =. An International Handbook , author =. 2009 , lastchecked =. doi:doi:10.1515/9783110213881.2.1212 , isbn =

  67. [76]

    Accurate Methods for the Statistics of Surprise and Coincidence

    Dunning, Ted. Accurate Methods for the Statistics of Surprise and Coincidence. Computational Linguistics. 1993

  68. [77]

    An in-depth look into the co-occurrence distribution of semantic associates , volume =

    Schulte Im Walde, Sabine and Melinger, Alissa , year =. An in-depth look into the co-occurrence distribution of semantic associates , volume =

  69. [78]

    Hypernyms under Siege: Linguistically-motivated Artillery for Hypernymy Detection

    Vered Shwartz and Enrico Santus and Dominik Schlechtweg. Hypernyms under Siege: Linguistically-motivated Artillery for Hypernymy Detection. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017

  70. [79]

    James A. Hampton. Typicality, Graded Membership, and Vagueness. Cognitive Science. 2007. doi:10.1080/15326900701326402

  71. [80]

    Linear Algebraic Structure of Word Senses, with Applications to Polysemy

    Arora, Sanjeev and Li, Yuanzhi and Liang, Yingyu and Ma, Tengyu and Risteski, Andrej. Linear Algebraic Structure of Word Senses, with Applications to Polysemy. Transactions of the Association for Computational Linguistics. 2018. doi:10.1162/tacl_a_00034

  72. [81]

    A New Formulation of Z ipf ' s Meaning-Frequency Law through Contextual Diversity

    Nagata, Ryo and Tanaka-Ishii, Kumiko. A New Formulation of Z ipf ' s Meaning-Frequency Law through Contextual Diversity. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. doi:10.18653/v1/2025.acl-long.744

  73. [82]

    Analysis and Evaluation of Language Models for Word Sense Disambiguation

    Loureiro, Daniel and Rezaee, Kiamehr and Pilehvar, Mohammad Taher and Camacho-Collados, Jose. Analysis and Evaluation of Language Models for Word Sense Disambiguation. Computational Linguistics. 2021. doi:10.1162/coli_a_00405

  74. [83]

    Towards Understanding Linear Word Analogies

    Ethayarajh, Kawin and Duvenaud, David and Hirst, Graeme. Towards Understanding Linear Word Analogies. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. doi:10.18653/v1/P19-1315

  75. [84]

    Identifying Linear Relational Concepts in Large Language Models

    Chanin, David and Hunter, Anthony and Camburu, Oana-Maria. Identifying Linear Relational Concepts in Large Language Models. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1:...

  76. [85]

    Are Large Language Models Good at Lexical Semantics? A Case of Taxonomy Learning

    Moskvoretskii, Viktor and Panchenko, Alexander and Nikishina, Irina. Are Large Language Models Good at Lexical Semantics? A Case of Taxonomy Learning. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-C...

  77. [86]

    Proceedings of the 40th International Conference on Machine Learning , articleno =

    Biderman, Stella and Schoelkopf, Hailey and Anthony, Quentin and Bradley, Herbie and O'Brien, Kyle and Hallahan, Eric and Khan, Mohammad Aflah and Purohit, Shivanshu and Prashanth, USVSN Sai and Raff, Edward and Skowron, Aviya and Sutawika, Lintang and Van Der Wal, Oskar , tit...

  78. [87]

    1975 , issn =

    Cognitive reference points , journal =. 1975 , issn =. doi:https://doi.org/10.1016/0010-0285(75)90021-3 , url =

  79. [88]

    Cognitive representations of semantic categories

    Eleanor Rosch. Cognitive representations of semantic categories. Journal of Experimental Psychology: General. 1975. doi:10.1037/0096-3445.104.3.192

  80. [89]

    Eleanor H. Rosch. Natural categories. Cognitive Psychology. 1973. doi:10.1016/0010-0285(73)90017-0

  81. [90]

    Antonymy and Canonicity: Experimental and Distributional Evidence

    Andreana Pastena and Alessandro Lenci. Antonymy and Canonicity: Experimental and Distributional Evidence. Proceedings of the 5th Workshop on Cognitive Aspects of the Lexicon (CogALex - V). 2016

  82. [91]

    Good and Bad Opposites: Using Textual and Experimental Techniques to Measure Antonym Canonicity

    Carita Paradis and Caroline Willners and Steven Jones. Good and Bad Opposites: Using Textual and Experimental Techniques to Measure Antonym Canonicity. The Mental Lexicon. 2009. doi:10.1075/ml.4.3.04par

  83. [92]

    Talking Heads: Understanding Inter-Layer Communication in Transformer Language Models , url =

    Merullo, Jack and Eickhoff, Carsten and Pavlick, Ellie , booktitle =. Talking Heads: Understanding Inter-Layer Communication in Transformer Language Models , url =. doi:10.52202/079017-1962 , editor =

  84. [93]

    and Leacock, Claudia and Tengi, Randee and Bunker, Ross T

    Miller, George A. and Leacock, Claudia and Tengi, Randee and Bunker, Ross T. A Semantic Concordance. H uman L anguage T echnology: Proceedings of a Workshop Held at Plainsboro, New Jersey, March 21-24, 1993. 1993

  85. [94]

    Psychological Methods , author =

    Generalized eta and omega squared statistics: measures of effect size for some common research designs , volume =. Psychological Methods , author =. 2003 , pmid =. doi:10.1037/1082-989X.8.4.434 , abstract =

  86. [95]

    Do Supervised Distributional Methods Really Learn Lexical Inference Relations?

    Levy, Omer and Remus, Steffen and Biemann, Chris and Dagan, Ido. Do Supervised Distributional Methods Really Learn Lexical Inference Relations?. Proceedings of the 2015 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language T...

  87. [96]

    Transparency Helps Reveal When Language Models Learn Meaning

    Wu, Zhaofeng and Merrill, William and Peng, Hao and Beltagy, Iz and Smith, Noah A. Transparency Helps Reveal When Language Models Learn Meaning. Transactions of the Association for Computational Linguistics. 2023. doi:10.1162/tacl_a_00565

  88. [97]

    What Does BERT Learn about the Structure of Language?

    Jawahar, Ganesh and Sagot, Beno \^i t and Seddah, Djam \'e. What Does BERT Learn about the Structure of Language?. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. doi:10.18653/v1/P19-1356

  89. [98]

    Linguistic Blind Spots of Large Language Models

    Cheng, Jiali and Amiri, Hadi. Linguistic Blind Spots of Large Language Models. Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics. 2025. doi:10.18653/v1/2025.cmcl-1.3

  90. [99]

    2024 , issue_date =

    Cao, Jiahang and Fang, Jinyuan and Meng, Zaiqiao and Liang, Shangsong , title =. 2024 , issue_date =. doi:10.1145/3643806 , journal =

  91. [100]

    Translating Embeddings for Modeling Multi-relational Data , url =

    Bordes, Antoine and Usunier, Nicolas and Garcia-Duran, Alberto and Weston, Jason and Yakhnenko, Oksana , booktitle =. Translating Embeddings for Modeling Multi-relational Data , url =

  92. [101]

    Proceedings of the International Conference on Learning Representations , year =

    RotatE: Knowledge Graph Embedding by Relational Rotation in Complex Space , author =. Proceedings of the International Conference on Learning Representations , year =

  93. [102]

    A Non-commutative Bilinear Model for Answering Path Queries in Knowledge Graphs

    Hayashi, Katsuhiko and Shimbo, Masashi. A Non-commutative Bilinear Model for Answering Path Queries in Knowledge Graphs. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Proces...

  94. [103]

    2022 , isbn =

    Yu, Jinxing and Cai, Yunfeng and Sun, Mingming and Li, Ping , title =. 2022 , isbn =. doi:10.1145/3511095.3531284 , booktitle =

  95. [104]

    Kilgarriff, Adam , editor =. How. Text,. 2004 , keywords =. doi:10.1007/978-3-540-30120-2_14 , abstract =

  96. [105]

    A Testset for Context-Aware LLM Translation in K orean-to- E nglish Discourse Level Translation

    Lee, Minjae and Noh, Youngbin and Lee, Seung Jin. A Testset for Context-Aware LLM Translation in K orean-to- E nglish Discourse Level Translation. Proceedings of the 31st International Conference on Computational Linguistics. 2025

  97. [106]

    Automatic Generation and Evaluation of Reading Comprehension Test Items with Large Language Models

    S. Automatic Generation and Evaluation of Reading Comprehension Test Items with Large Language Models. Proceedings of the 3rd Workshop on Tools and Resources for People with REAding DIfficulties (READI) @ LREC-COLING 2024. 2024

  98. [107]

    Do Large Language Models Understand Word Senses?

    Meconi, Domenico and Stirpe, Simone and Martelli, Federico and Lavalle, Leonardo and Navigli, Roberto. Do Large Language Models Understand Word Senses?. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18653/v1/2025.emnlp-main.1720

  99. [108]

    Testing Z ipf ' s meaning-frequency law with wordnets as sense inventories

    Bond, Francis and Janz, Arkadiusz and Maziarz, Marek and Rudnicka, Ewa. Testing Z ipf ' s meaning-frequency law with wordnets as sense inventories. Proceedings of the 10th Global Wordnet Conference. 2019. doi:10.18653/v1/2019.gwc-1.44

  100. [109]

    Does Large Language Model Contain Task-Specific Neurons?

    Song, Ran and He, Shizhu and Jiang, Shuting and Xian, Yantuan and Gao, Shengxiang and Liu, Kang and Yu, Zhengtao. Does Large Language Model Contain Task-Specific Neurons?. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.1865...

  101. [110]

    Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT

    Wu, Shijie and Dredze, Mark. Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)....

  102. [111]

    Can Large Language Models Robustly Perform Natural Language Inference for J apanese Comparatives?

    Mikami, Yosuke and Matsuoka, Daiki and Yanaka, Hitomi. Can Large Language Models Robustly Perform Natural Language Inference for J apanese Comparatives?. Proceedings of the 16th International Conference on Computational Semantics. 2025

  103. [112]

    Proceedings of the 37th International Conference on Neural Information Processing Systems , articleno =

    Valeriani, Lucrezia and Doimo, Diego and Cuturello, Francesca and Laio, Alessandro and Ansuini, Alessio and Cazzaniga, Alberto , title =. Proceedings of the 37th International Conference on Neural Information Processing Systems , articleno =. 2023 , publisher =

  104. [113]

    Exploring L atin W ord N et synset annotation with LLM s

    Santoro, Daniela and Marchesi, Beatrice and Zampetta, Silvia and Tredici, Marco Del and Biagetti, Erica and Litta, Eleonora and Combei, Claudia Roberta and Rocchi, Stefano and Facchinetti, Tullio and Ginevra, Riccardo and Zanchi, Chiara. Exploring L atin W ord N et synset anno...

  105. [114]

    Leveraging LLM s for Constructing W ord N ets Automatically as Bilingual Resources

    Bergh, Johann and Waitelonis, J. Leveraging LLM s for Constructing W ord N ets Automatically as Bilingual Resources. Proceedings of the 13th Global Wordnet Conference. 2025. doi:10.18653/v1/2025.gwc-1.24

  106. [115]

    Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence,

    Recent Trends in Word Sense Disambiguation: A Survey , author =. Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence,. 2021 , month =. doi:10.24963/ijcai.2021/593 , url =

  107. [116]

    User's guide to correlation coefficients , year =

    Akoglu, Haldun , title =. User's guide to correlation coefficients , year =. doi:10.1016/j.tjem.2018.08.001 , volume =

  108. [117]

    1946 , lastchecked =

    Mathematical Methods of Statistics , author =. 1946 , lastchecked =. doi:doi:10.1515/9781400883868 , isbn =

  109. [118]

    Kotrl , year=

    Joe W. Kotrl , year=. Information Technology, Learning, and Performance Journal , title=

  110. [119]

    Fuglede and F

    B. Fuglede and F. Topsoe , doi =. Jensen-Shannon divergence and Hilbert space embedding , url =. International Symposium on Information Theory 2004 Proceedings. , year =

  111. [120]

    Proceedings of the Association for Information Science and Technology , author =

    Paradigmatic relations and syntagmatic relations:. Proceedings of the Association for Information Science and Technology , author =

  112. [121]

    Journal of Language Modelling , author =

    Distinguishing between paradigmatic semantic relations across word classes: human ratings and distributional similarity , volume =. Journal of Language Modelling , author =. 2020 , keywords =. doi:10.15398/jlm.v8i1.199 , abstract =

  113. [122]

    Contrasting

    Lapesa, Gabriella and Evert, Stefan and Schulte im Walde, Sabine , editor =. Contrasting. Proceedings of the. 2014 , keywords =. doi:10.3115/v1/S14-1020 , urldate =

  114. [123]

    Lynne Murphy and Caroline Willners

    Steven Jones and Carita Paradis and M. Lynne Murphy and Caroline Willners. Googling for ‘Opposites’: a Web-based Study of Antonym Canonicity. Corpora. 2007. doi:10.3366/cor.2007.2.2.129

  115. [124]

    Lynne Murphy

    M. Lynne Murphy. Semantic Relations and the Lexicon: Antonymy, Synonymy and other Paradigms. 2003. doi:10.1017/CBO9780511486494

  116. [125]

    The Semantic Relations in LLM s: An Information-theoretic Compression Approach

    Tseng, Yu-Hsiang and Chen, Pin-Er and Lian, Da-Chen and Hsieh, Shu-Kai. The Semantic Relations in LLM s: An Information-theoretic Compression Approach. Proceedings of the Workshop: Bridging Neurons and Symbols for Natural Language Processing and Knowledge Graphs Reasoning (Neu...

  117. [126]

    The Llama 3 Herd of Models , journal =

    Abhimanyu Dubey and Abhinav Jauhri and Abhinav Pandey and Abhishek Kadian and Ahmad Al. The Llama 3 Herd of Models , journal =. 2024 , url =. doi:10.48550/ARXIV.2407.21783 , eprinttype =. 2407.21783 , timestamp =

  118. [127]

    Similarity of Semantic Relations

    Turney, Peter D. Similarity of Semantic Relations. Computational Linguistics. 2006. doi:10.1162/coli.2006.32.3.379

  119. [128]

    S im L ex-999: Evaluating Semantic Models With (Genuine) Similarity Estimation

    Hill, Felix and Reichart, Roi and Korhonen, Anna. S im L ex-999: Evaluating Semantic Models With (Genuine) Similarity Estimation. Computational Linguistics. 2015. doi:10.1162/COLI_a_00237

  120. [129]

    Proceedings of the 14th International Joint Conference on Artificial Intelligence - Volume 1 , pages =

    Resnik, Philip , title =. Proceedings of the 14th International Joint Conference on Artificial Intelligence - Volume 1 , pages =. 1995 , isbn =

  121. [130]

    Evaluating W ord N et-based Measures of Lexical Semantic Relatedness

    Budanitsky, Alexander and Hirst, Graeme. Evaluating W ord N et-based Measures of Lexical Semantic Relatedness. Computational Linguistics. 2006. doi:10.1162/coli.2006.32.1.13

  122. [131]

    Princeton WordNet Gloss Corpus , year =

  123. [132]

    A Review of Relational Machine Learning for Knowledge Graphs , year=

    Nickel, Maximilian and Murphy, Kevin and Tresp, Volker and Gabrilovich, Evgeniy , journal=. A Review of Relational Machine Learning for Knowledge Graphs , year=

  124. [133]

    1995 , doi =

    Bland, J Martin and Altman, Douglas G , title =. 1995 , doi =. https://www.bmj.com/content/310/6973/170.full.pdf , journal =

  125. [134]

    Anomalies in the W ord N et Verb Hierarchy

    Richens, Tom. Anomalies in the W ord N et Verb Hierarchy. Proceedings of the 22nd International Conference on Computational Linguistics (Coling 2008). 2008

  126. [135]

    Some structural tests for W ord N et with results

    Lohk, Ahti and Orav, Heili and V \ o handu, Leo. Some structural tests for W ord N et with results. Proceedings of the Seventh Global W ordnet Conference. 2014. doi:10.18653/v1/W14-0143

  127. [136]

    Meta-Learning a Cross-lingual Manifold for Semantic Parsing

    Sherborne, Tom and Lapata, Mirella. Meta-Learning a Cross-lingual Manifold for Semantic Parsing. Transactions of the Association for Computational Linguistics. 2023. doi:10.1162/tacl_a_00533

  128. [137]

    Calibrated Interpretation: Confidence Estimation in Semantic Parsing

    Stengel-Eskin, Elias and Van Durme, Benjamin. Calibrated Interpretation: Confidence Estimation in Semantic Parsing. Transactions of the Association for Computational Linguistics. 2023. doi:10.1162/tacl_a_00598

  129. [138]

    Conditional Semantic Textual Similarity via Conditional Contrastive Learning

    Liu, Xinyue and Qin, Zeyang and Wang, Zeyu and Liang, Wenxin and Zong, Linlin and Xu, Bo. Conditional Semantic Textual Similarity via Conditional Contrastive Learning. Proceedings of the 31st International Conference on Computational Linguistics. 2025

  130. [139]

    Rethinking Word Similarity: Semantic Similarity through Classification Confusion

    Zhou, Kaitlyn and Gao, Haishan and Chen, Sarah Li and Edelstein, Dan and Jurafsky, Dan and Shani, Chen. Rethinking Word Similarity: Semantic Similarity through Classification Confusion. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Associatio...

  131. [140]

    Current Semantic-change Quantification Methods Struggle with Discovery in the Wild

    Umarova, Khonzoda and Lee, Lillian and Kim, Laerdon. Current Semantic-change Quantification Methods Struggle with Discovery in the Wild. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18653/v1/2025.emnlp-main.1791

  132. [141]

    Analyzing Semantic Change through Lexical Replacements

    Periti, Francesco and Cassotti, Pierluigi and Dubossarsky, Haim and Tahmasebi, Nina. Analyzing Semantic Change through Lexical Replacements. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. doi:10.18653/v1/2...

  133. [142]

    Semantic Change Characterization with LLM s using Rhetorics

    de S \'a , J \'a der Martins Camboim and Lee, Jooyoung and Da Silveira, Marcos and Pruski, Cedric. Semantic Change Characterization with LLM s using Rhetorics. The Proceedings for the 6th International Workshop on Computational Approaches to Language Change ( LC hange ' 26). 2...

  134. [143]

    Proceedings of the AAAI Conference on Artificial Intelligence , author=

    Is Word Sense Disambiguation Dead in the LLM Era? , volume=. Proceedings of the AAAI Conference on Artificial Intelligence , author=. 2026 , month=. doi:10.1609/aaai.v40i46.41331 , number=

  135. [144]

    In the LLM era, Word Sense Induction remains unsolved

    Mosolova, Anna and Candito, Marie and Ramisch, Carlos. In the LLM era, Word Sense Induction remains unsolved. Findings of the Association for Computational Linguistics: ACL 2025. 2025. doi:10.18653/v1/2025.findings-acl.882

  136. [145]

    Proceedings of the 41st International Conference on Machine Learning , articleno =

    Park, Kiho and Choe, Yo Joong and Veitch, Victor , title =. Proceedings of the 41st International Conference on Machine Learning , articleno =. 2024 , publisher =

  137. [146]

    Kingma and Jimmy Ba , title=

    Diederik P. Kingma and Jimmy Ba , title=. 3rd International Conference on Learning Representations , year=

  138. [147]

    The Twelfth International Conference on Learning Representations,

    Lukas Berglund and Meg Tong and Maximilian Kaufmann and Mikita Balesni and Asa Cooper Stickland and Tomasz Korbak and Owain Evans , title =. The Twelfth International Conference on Learning Representations,. 2024 , url =

  139. [148]

    Slate , year=

    Early Findings in Using LLMs to Assess Semantic Relations Strength (Short Paper) , author=. Slate , year=

  140. [149]

    The Twelfth International Conference on Learning Representations,

    Evan Hernandez and Arnab Sen Sharma and Tal Haklay and Kevin Meng and Martin Wattenberg and Jacob Andreas and Yonatan Belinkov and David Bau , title =. The Twelfth International Conference on Learning Representations,. 2024 , url =

  141. [150]

    CoRR , volume =

    Vishwas Mruthyunjaya and Pouya Pezeshkpour and Estevam Hruschka and Nikita Bhutani , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2308.13676 , eprinttype =. 2308.13676 , timestamp =

  142. [151]

    Unraveling Antonym ' s Word Vectors through a S iamese-like Network

    Etcheverry, Mathias and Wonsever, Dina. Unraveling Antonym ' s Word Vectors through a S iamese-like Network. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. doi:10.18653/v1/P19-1319

  143. [152]

    and Dorr, Bonnie J

    Mohammad, Saif M. and Dorr, Bonnie J. and Hirst, Graeme and Turney, Peter D. Computing Lexical Contrast. Computational Linguistics. 2013. doi:10.1162/COLI_a_00143

  144. [153]

    Discriminating between Lexico-Semantic Relations with the Specialization Tensor Model

    Glava s , Goran and Vuli \'c , Ivan. Discriminating between Lexico-Semantic Relations with the Specialization Tensor Model. Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2...

  145. [154]

    Word Embedding-based Antonym Detection using Thesauri and Distributional Information

    Ono, Masataka and Miwa, Makoto and Sasaki, Yutaka. Word Embedding-based Antonym Detection using Thesauri and Distributional Information. Proceedings of the 2015 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technolog...

  146. [155]

    A Mixture-of-Experts Model for Antonym-Synonym Discrimination

    Zhipeng Xie and Nan Zeng. A Mixture-of-Experts Model for Antonym-Synonym Discrimination. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)....

  147. [156]

    Antonym-Synonym Classification Based on New Sub-Space Embeddings

    Muhammad Asif Ali and Yifang Sun and Xiaoling Zhou and Wei Wang and Xiang Zhao. Antonym-Synonym Classification Based on New Sub-Space Embeddings. Proceedings of the AAAI Conference on Artificial Intelligence. 2019. doi:10.1609/AAAI.V33I01.33016204

  148. [157]

    Learning to Distinguish Hypernyms and Co-Hyponyms

    Julie Weeds and Daoud Clarke and Jeremy Reffin and David Weir and Bill Keller. Learning to Distinguish Hypernyms and Co-Hyponyms. Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers. 2014

  149. [158]

    Uncovering Distributional Differences between Synonyms and Antonyms in a Word Space Model

    Silke Scheible and Sabine Schulte Im Walde and Sylvia Springorum. Uncovering Distributional Differences between Synonyms and Antonyms in a Word Space Model. Proceedings of the Sixth International Joint Conference on Natural Language Processing. 2013

  150. [159]

    Collective nouns, aggregate nouns, and superordinates

    Frank Joosten. Collective nouns, aggregate nouns, and superordinates. Lingvisticae Investigationes. 2010. doi:10.1075/li.33.1.03joo

  151. [160]

    A comparison of hyponym and synonym decisions

    Roger Chaffin and Arnold Glass. A comparison of hyponym and synonym decisions. Journal of Psycholinguistic Research. 1990. doi:10.1007/BF01077260

  152. [161]

    The similarity and diversity of semantic relations

    Roger Chaffin and H H Clark. The similarity and diversity of semantic relations. Memory & Cognition. 1984

  153. [162]

    Llama 2: Open Foundation and Fine-Tuned Chat Models

    Hugo Touvron and Louis Martin and Kevin Stone and Peter Albert and Amjad Almahairi and Yasmine Babaei and Nikolay Bashlykov and Soumya Batra and Prajjwal Bhargava and Shruti Bhosale and Dan Bikel and Lukas Blecher and Cristian Canton - Ferrer and Moya Chen and Guillem Cucurull...

  154. [163]

    Susan Zhang and Stephen Roller and Naman Goyal and Mikel Artetxe and Moya Chen and Shuohui Chen and Christopher Dewan and Mona T. Diab and Xian Li and Xi Victoria Lin and Todor Mihaylov and Myle Ott and Sam Shleifer and Kurt Shuster and Daniel Simig and Punit Singh Koura and A...

  155. [164]

    Battig and William E

    William F. Battig and William E. Montague. Category norms of verbal items in 56 categories A replication and extension of the Connecticut category norms. Journal of Experimental Psychology. 1969. doi:10.1037/h0027577

  156. [165]

    Nelson and Cathy L

    Douglas L. Nelson and Cathy L. McEvoy and Thomas A. Schreiber. The University of South Florida free association, rhyme, and word fragment norms. Behavior Research Methods, Instruments, and Computers. 2004. doi:10.3758/BF03195588/METRICS

  157. [166]

    Category norms: An updated and expanded version of the Battig and Montague (1969) norms

    James P Van Overschelde and Katherine A Rawson and John Dunlosky. Category norms: An updated and expanded version of the Battig and Montague (1969) norms. Journal of Memory and Language. 2004. doi:10.1016/j.jml.2003.10.003

  158. [167]

    Acquisition of Turkish meronym based on classification of patterns

    Tuǧba Yıldız and Banu Diri and Savaş Yıldırım. Acquisition of Turkish meronym based on classification of patterns. Pattern Analysis and Applications. 2016. doi:10.1007/S10044-015-0516-9/TABLES/13

  159. [168]

    Mining Text Patterns for Synonyms Extraction

    Andrey Simanovsky and Alexander Ulanov. Mining Text Patterns for Synonyms Extraction. 2011 22nd International Workshop on Database and Expert Systems Applications. 2011. doi:10.1109/DEXA.2011.53

  160. [169]

    Learning Semantic Network Patterns for Hypernymy Extraction

    Tim Vor Der Brück. Learning Semantic Network Patterns for Hypernymy Extraction. Proceedings of the 6th Workshop on Ontologies and Lexical Resources. 2010

  161. [170]

    Judea Pearl and Madelyn Glamour and Nicholas P. Jewell. Causal Inference in Statistics: A Primer. 2016

  162. [171]

    Mann `` is to ``Donna'' as「国王」is to Reine Adapting the Analogy Task for Multilingual and Contextual Embeddings

    Mickus, Timothee and Cal \`o , Eduardo and Jacqmin, L \'e o and Paperno, Denis and Constant, Mathieu. Mann `` is to ``Donna'' as「国王」is to Reine Adapting the Analogy Task for Multilingual and Contextual Embeddings. Proceedings of the 12th Joint Conference on Lexical and Computa...

  163. [172]

    Causality: Models, Reasoning and Inference

    Judea Pearl. Causality: Models, Reasoning and Inference. 2009

  164. [173]

    Semantics

    John I Saeed. Semantics. 2015

  165. [174]

    McNamara

    Timothy P. McNamara. Semantic Priming. 2005. doi:10.4324/9780203338001

  166. [175]

    Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books

    Yukun Zhu and Ryan Kiros and Rich Zemel and Ruslan Salakhutdinov and Raquel Urtasun and Antonio Torralba and Sanja Fidler. Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books. 2015 IEEE International Conference on Computer Vis...

  167. [176]

    Thomas McCoy and Najoung Kim and Benjamin Van Durme and Samuel R

    Ian Tenney and Patrick Xia and Berlin Chen and Alex Wang and Adam Poliak and R. Thomas McCoy and Najoung Kim and Benjamin Van Durme and Samuel R. Bowman and Dipanjan Das and Ellie Pavlick. What do you learn from context? Probing for sentence structure in contextualized word re...

  168. [177]

    Bird's eye: Probing for linguistic graph structures with a simple information-theoretic approach

    Yifan Hou and Mrinmaya Sachan. Bird's eye: Probing for linguistic graph structures with a simple information-theoretic approach. ACL-IJCNLP 2021 - 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Langua...

  169. [178]

    Post-hoc Interpretability for Neural NLP: A Survey

    Andreas Madsen and Siva Reddy and Sarath Chandar. Post-hoc Interpretability for Neural NLP: A Survey. ACM Computing Surveys. 2021. doi:10.1145/inreview

  170. [179]

    Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing

    Pengfei Liu and Weizhe Yuan and Jinlan Fu and Zhengbao Jiang and Hiroaki Hayashi and Graham Neubig. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Computing Surveys. 2022. doi:10.1145/3560815

  171. [180]

    ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

    Zhenzhong Lan and Mingda Chen and Sebastian Goodman and Kevin Gimpel and Piyush Sharma and Radu Soricut. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, Ap...

  172. [181]

    Should You Mask 15 \

    Alexander Wettig and Tianyu Gao and Zexuan Zhong and Danqi Chen. Should You Mask 15 \. CoRR. 2022. 2202.08005

  173. [182]

    Proceedings of the AAAI Conference on Artificial Intelligence , author=

    KEML: A Knowledge-Enriched Meta-Learning Framework for Lexical Relation Classification , volume=. Proceedings of the AAAI Conference on Artificial Intelligence , author=. 2021 , month=. doi:10.1609/aaai.v35i15.17640 , number=

  174. [183]

    RoBERTa: A Robustly Optimized BERT Pretraining Approach

    Yinhan Liu and Myle Ott and Naman Goyal and Jingfei Du and Mandar Joshi and Danqi Chen and Omer Levy and Mike Lewis and Luke Zettlemoyer and Veselin Stoyanov. RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR. 2019. 1907.11692

  175. [184]

    DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

    Victor Sanh and Lysandre Debut and Julien Chaumond and Thomas Wolf. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. CoRR. 2019. 1910.01108

  176. [185]

    Automatic Acquisition of Hyponyms from Large Text Corpora

    Marti A Hearst. Automatic Acquisition of Hyponyms from Large Text Corpora. COLING 1992 Volume 2: The 14th International Conference on Computational Linguistics. 1992

  177. [186]

    George A. Miller. WordNet. Communications of the ACM. 1995. doi:10.1145/219717.219748

  178. [187]

    and Childers, Donald G

    Fischler, Ira and Bloom, Paul A. and Childers, Donald G. and Roucos, Salim E. and Perry Jr., Nathan W. , title =. Psychophysiology , volume =. doi:https://doi.org/10.1111/j.1469-8986.1983.tb00920.x , url =. https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1469-8986.1983.tb009...

  179. [188]

    Causalm: Causal model explanation through counterfactual language models

    Amir Feder and Nadav Oved and Uri Shalit and Roi Reichart. Causalm: Causal model explanation through counterfactual language models. Computational Linguistics. 2021. doi:10.1162/COLI_a_00404

  180. [189]

    Brown and Alan B

    Morton B. Brown and Alan B. Forsythe , doi =. Robust Tests for the Equality of Variances , volume =. Journal of the American Statistical Association , month =

  181. [190]

    HyperLex: A Large-Scale Evaluation of Graded Lexical Entailment

    Ivan Vulić and Daniela Gerz and Douwe Kiela and Felix Hill and Anna Korhonen. HyperLex: A Large-Scale Evaluation of Graded Lexical Entailment. Computational Linguistics. 2017. doi:10.1162/COLI_a_00301

  182. [191]

    Automatic Discovery of Part-Whole Relations

    Roxana Girju and Dan Moldovan and Adriana Badulescu. Automatic Discovery of Part-Whole Relations. Computational Linguistics. 2006. doi:10.1162/COLI.2006.32.1.83

  183. [192]

    Information-theoretic probing with minimum description length

    Elena Voita and Ivan Titov. Information-theoretic probing with minimum description length. EMNLP 2020 - 2020 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference. 2020. doi:10.18653/v1/2020.emnlp-main.14

  184. [193]

    Designing and interpreting probes with control tasks

    John Hewitt and Percy Liang. Designing and interpreting probes with control tasks. EMNLP-IJCNLP 2019 - 2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing, Proceedings of the Conference. 2020...

  185. [194]

    Do language embeddings capture scales?

    Xikun Zhang and Deepak Ramachandran and Ian Tenney and Yanai Elazar and Dan Roth. Do language embeddings capture scales?. Findings of the Association for Computational Linguistics Findings of ACL: EMNLP 2020. 2020. doi:10.18653/v1/2020.findings-emnlp.439

  186. [195]

    Exploring BERT’s Sensitivity to Lexical Cues using Tests from Semantic Priming

    Kanishka Misra and Allyson Ettinger and Julia Taylor Rayz. Exploring BERT’s Sensitivity to Lexical Cues using Tests from Semantic Priming. Findings of the Association for Computational Linguistics Findings of ACL: EMNLP 2020. 2020. doi:10.18653/V1/2020.FINDINGS-EMNLP.415

  187. [196]

    How Pre-trained Language Models Capture Factual Knowledge? A Causal-Inspired Analysis

    Shaobo Li and Xiaoguang Li and Lifeng Shang and Zhenhua Dong and Chengjie Sun and Bingquan Liu and Zhenzhou Ji and Xin Jiang and Qun Liu. How Pre-trained Language Models Capture Factual Knowledge? A Causal-Inspired Analysis. Findings of the Association for Computational Lingui...

  188. [197]

    Explaining NLP Models via Minimal Contrastive Editing (MiCE)

    Alexis Ross and Ana Marasović and Matthew Peters. Explaining NLP Models via Minimal Contrastive Editing (MiCE). Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. 2021. doi:10.18653/v1/2021.findings-acl.336

  189. [198]

    John Hewitt and Christopher D. Manning. A structural probe for finding syntax in word representations. NAACL HLT 2019 - 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference. ...

  190. [199]

    Language Models are Unsupervised Multitask Learners

    Alec Radford and Jeffrey Wu and Rewon Child and David Luan and Dario Amodei and Ilya Sutskever. Language Models are Unsupervised Multitask Learners. OpenAI. 2019

  191. [200]

    Mihalcea and Dan I

    Rada F. Mihalcea and Dan I. Moldovan. Pattern Learning and Active Feature Selection for Word Sense Disambiguation. Proceedings of \ SENSEVAL\ -2 Second International Workshop on Evaluating Word Sense Disambiguation Systems. 2001

  192. [201]

    A Shortest Path Dependency Kernel for Relation Extraction

    Razvan C Bunescu and Raymond J Mooney. A Shortest Path Dependency Kernel for Relation Extraction. Proceedings of Human Language Technology Conference and Conference on Empirical Methods in Natural Language Processing. 2005

  193. [202]

    Handling the Impact of Low Frequency Events on Co-occurrence based Measures of Word Similarity - A Case Study of Pointwise Mutual Information

    François Role and Mohamed Nadif. Handling the Impact of Low Frequency Events on Co-occurrence based Measures of Word Similarity - A Case Study of Pointwise Mutual Information. Proceedings of International Conference on Knowledge Discovery and Information Retrieval (KDIR-2011). 2011

  194. [203]

    Hypernym-LIBre: A free Web-based Corpus for Hypernym Detection

    Shaurya Rawat and Mariano Rico and Oscar Corcho. Hypernym-LIBre: A free Web-based Corpus for Hypernym Detection. Proceedings of the 12th Web as Corpus Workshop. 2020. doi:10.5281/zenodo.3662204

  195. [204]

    Entailment above the word level in distributional semantics

    Marco Baroni and Raffaella Bernardi and Ngoc-Quynh Do and Chung-Chieh Shan. Entailment above the word level in distributional semantics. Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics. 2012

  196. [205]

    Distinguishing Antonyms and Synonyms in a Pattern-based Neural Network

    Kim Anh Nguyen and Sabine Schulte Im Walde and Ngoc Thang Vu. Distinguishing Antonyms and Synonyms in a Pattern-based Neural Network. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017

  197. [206]

    PATTY: A Taxonomy of Relational Patterns with Semantic Types

    Ndapandula Nakashole and Gerhard Weikum and Fabian Suchanek. PATTY: A Taxonomy of Relational Patterns with Semantic Types. Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning. 2012

  198. [207]

    Representing Text for Joint Embedding of Text and Knowledge Bases

    Kristina Toutanova and Danqi Chen and Patrick Pantel and Hoifung Poon and Pallavi Choudhury and Michael Gamon. Representing Text for Joint Embedding of Text and Knowledge Bases. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. doi:1...

  199. [208]

    GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

    Alex Wang and Amanpreet Singh and Julian Michael and Felix Hill and Omer Levy and Samuel Bowman. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Network...

  200. [209]

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

    Jacob Devlin and Ming-Wei Chang and Kenton Lee and Kristina Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lan...

  201. [210]

    Language Models as Knowledge Bases?

    Fabio Petroni and Tim Rocktäschel and Sebastian Riedel and Patrick Lewis and Anton Bakhtin and Yuxiang Wu and Alexander Miller. Language Models as Knowledge Bases?. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International...

  202. [211]

    Transformers: State-of-the-Art Natural Language Processing

    Thomas Wolf and Lysandre Debut and Victor Sanh and Julien Chaumond and Clement Delangue and Anthony Moi and Pierric Cistac and Tim Rault and Remi Louf and Morgan Funtowicz and Joe Davison and Sam Shleifer and Patrick von Platen and Clara Ma and Yacine Jernite and Julien Plu an...

  203. [212]

    Putting Words in BERT’s Mouth: Navigating Contextualized Vector Spaces with Pseudowords

    Taelin Karidi and Yichu Zhou and Nathan Schneider and Omri Abend and Vivek Srikumar. Putting Words in BERT’s Mouth: Navigating Contextualized Vector Spaces with Pseudowords. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. doi:10.18...

  204. [213]

    Exploring the Role of BERT Token Representations to Explain Sentence Probing Results

    Hosein Mohebbi and Ali Modarressi and Mohammad Taher Pilehvar. Exploring the Role of BERT Token Representations to Explain Sentence Probing Results. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. doi:10.18653/v1/2021.emnlp-main.61

  205. [214]

    Espresso: Leveraging Generic Patterns for Automatically Harvesting Semantic Relations

    Patrick Pantel and Marco Pennacchiotti. Espresso: Leveraging Generic Patterns for Automatically Harvesting Semantic Relations. Proceedings of the 21st International Conference on Computational Linguistics and the 44th annual meeting of the ACL - ACL '06. 2006. doi:10.3115/1220...

  206. [215]

    Pragmatic competence of pre-trained language models through the lens of discourse connectives

    Lalchand Pandia and Yan Cong and Allyson Ettinger. Pragmatic competence of pre-trained language models through the lens of discourse connectives. Proceedings of the 25th Conference on Computational Natural Language Learning. 2021. doi:10.18653/v1/2021.conll-1.29

  207. [216]

    Do Neural Language Models Overcome Reporting Bias?

    Vered Shwartz and Yejin Choi. Do Neural Language Models Overcome Reporting Bias?. Proceedings of the 28th International Conference on Computational Linguistics. 2020. doi:10.18653/v1/2020.coling-main.605

  208. [217]

    EVALution 1.0: an Evolving Semantic Dataset for Training and Evaluation of Distributional Semantic Models

    Enrico Santus and Frances Yung and Alessandro Lenci and Chu-Ren Huang. EVALution 1.0: an Evolving Semantic Dataset for Training and Evaluation of Distributional Semantic Models. Proceedings of the 4th Workshop on Linked Data in Linguistics: Resources and Applications. 2015. do...

  209. [218]

    Exploiting Image Generality for Lexical Entailment Detection

    Douwe Kiela and Laura Rimell and Ivan Vulić and Stephen Clark. Exploiting Image Generality for Lexical Entailment Detection. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language P...

  210. [219]

    Improving Hypernymy Detection with an Integrated Path-based and Distributional Method

    Vered Shwartz and Yoav Goldberg and Ido Dagan. Improving Hypernymy Detection with an Integrated Path-based and Distributional Method. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. doi:10.18653/v1/P16-1226

  211. [220]

    Semantically Equivalent Adversarial Rules for Debugging NLP models

    Marco Tulio Ribeiro and Sameer Singh and Carlos Guestrin. Semantically Equivalent Adversarial Rules for Debugging NLP models. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. doi:10.18653/v1/P18-1079

  212. [221]

    Hearst Patterns Revisited: Automatic Hypernym Detection from Large Text Corpora

    Stephen Roller and Douwe Kiela and Maximilian Nickel. Hearst Patterns Revisited: Automatic Hypernym Detection from Large Text Corpora. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2018. doi:10.18653/v1/P18-2057

  213. [222]

    BERT Rediscovers the Classical NLP Pipeline

    Ian Tenney and Dipanjan Das and Ellie Pavlick. BERT Rediscovers the Classical NLP Pipeline. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. doi:10.18653/v1/P19-1452

  214. [223]

    A Tale of a Probe and a Parser

    Rowan Hall Maudslay and Josef Valvoda and Tiago Pimentel and Adina Williams and Ryan Cotterell. A Tale of a Probe and a Parser. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.659

  215. [224]

    Knowledgeable or Educated Guess? Revisiting Language Models as Knowledge Bases

    Boxi Cao and Hongyu Lin and Xianpei Han and Le Sun and Lingyong Yan and Meng Liao and Tong Xue and Jin Xu. Knowledgeable or Educated Guess? Revisiting Language Models as Knowledge Bases. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics an...

  216. [225]

    Knowledge Neurons in Pretrained Transformers

    Damai Dai and Li Dong and Yaru Hao and Zhifang Sui and Baobao Chang and Furu Wei. Knowledge Neurons in Pretrained Transformers. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. doi:10.18653/v1/2022.acl-long.581

  217. [226]

    Can Prompt Probe Pretrained Language Models? Understanding the Invisible Risks from a Causal View

    Boxi Cao and Hongyu Lin and Xianpei Han and Fangchao Liu and Le Sun. Can Prompt Probe Pretrained Language Models? Understanding the Invisible Risks from a Causal View. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Paper...

  218. [227]

    Analyzing BERT’s Knowledge of Hypernymy via Prompting

    Michael Hanna and David Mareček. Analyzing BERT’s Knowledge of Hypernymy via Prompting. Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP. 2021. doi:10.18653/v1/2021.blackboxnlp-1.20

  219. [228]

    How we BLESSed distributional semantic evaluation

    Marco Baroni and Alessandro Lenci. How we BLESSed distributional semantic evaluation. Proceedings of the GEMS 2011 Workshop on GEometrical Models of Natural Language Semantics. 2011

  220. [229]

    On the Systematicity of Probing Contextualized Word Representations: The Case of Hypernymy in BERT

    Abhilasha Ravichander and Eduard Hovy and Kaheer Suleman and Adam Trischler and Jackie Chi Kit Cheung. On the Systematicity of Probing Contextualized Word Representations: The Case of Hypernymy in BERT. Proceedings of the Ninth Joint Conference on Lexical and Computational Sem...

  221. [230]

    Minimally Supervised Learning of Semantic Knowledge from Query Logs

    Mamoru Komachi and Hisami Suzuki. Minimally Supervised Learning of Semantic Knowledge from Query Logs. Proceedings of the Third International Joint Conference on Natural Language Processing: Volume-I. 2008

  222. [231]

    Lexical Patterns or Dependency Patterns: Which Is Better for Hypernym Extraction?

    Erik Tjong and Kim Sang and Katja Hofmann. Lexical Patterns or Dependency Patterns: Which Is Better for Hypernym Extraction?. Proceedings of the Thirteenth Conference on Computational Natural Language Learning (CoNLL-2009). 2009

  223. [232]

    UMBC EBIQUITY-CORE: Semantic Textual Similarity Systems

    Lushan Han and Abhay Kashyap and Tim Finin and James Mayfield and Jonathan Weese. UMBC EBIQUITY-CORE: Semantic Textual Similarity Systems. Second Joint Conference on Lexical and Computational Semantics (*SEM), Volume 1: Proceedings of the Main Conference and the Shared Task: S...

  224. [233]

    What bert is not: Lessons from a new suite of psycholinguistic diagnostics for language models

    Allyson Ettinger. What bert is not: Lessons from a new suite of psycholinguistic diagnostics for language models. Transactions of the Association for Computational Linguistics. 2020. doi:10.1162/TACL_A_00298/43535/WHAT-BERT-IS-NOT-LESSONS-FROM-A-NEW-SUITE-OF

  225. [234]

    oLMpics-On What Language Model Pre-training Captures

    Alon Talmor and Yanai Elazar and Yoav Goldberg and Jonathan Berant. oLMpics-On What Language Model Pre-training Captures. Transactions of the Association for Computational Linguistics. 2020. doi:10.1162/tacl_a_00342

  226. [235]

    Weld and Luke Zettlemoyer and Omer Levy

    Mandar Joshi and Danqi Chen and Yinhan Liu and Daniel S. Weld and Luke Zettlemoyer and Omer Levy. SpanBERT: Improving Pre-training by Representing and Predicting Spans. Transactions of the Association for Computational Linguistics. 2020. doi:10.1162/tacl_a_00300

  227. [236]

    Computing Word-Pair Antonymy

    Mohammad, Saif and Dorr, Bonnie and Hirst, Graeme. Computing Word-Pair Antonymy. Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing. 2008

  228. [237]

    Brno studies in English , author =

    The distributional asymmetries of. Brno studies in English , author =. 2017 , keywords =. doi:10.5817/BSE2017-1-1 , language =

  229. [238]

    Jones, Steven and Murphy, M. Lynne. Using corpora to investigate antonym acquisition. International Journal of Corpus Linguistics. 2005. doi:https://doi.org/10.1075/ijcl.10.3.06jon

  230. [239]

    Text & Talk , author =

    A lexico-syntactic analysis of antonym co-occurrence in spoken. Text & Talk , author =. 2006 , pages =. doi:doi:10.1515/TEXT.2006.009 , number =

  231. [240]

    Measuring and Improving Consistency in Pretrained Language Models

    Yanai Elazar and Nora Kassner and Shauli Ravfogel and Abhilasha Ravichander and Eduard Hovy and Hinrich Schütze and Yoav Goldberg. Measuring and Improving Consistency in Pretrained Language Models. Transactions of the Association for Computational Linguistics. 2021. doi:10.116...

  232. [241]

    Linguistic Regularities in Continuous Space Word Representations

    Mikolov, Tomas and Yih, Wen-tau and Zweig, Geoffrey. Linguistic Regularities in Continuous Space Word Representations. Proceedings of the 2013 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies. 2013

  233. [242]

    Word Embeddings, Analogies, and Machine Learning: Beyond king - man + woman = queen

    Drozd, Aleksandr and Gladkova, Anna and Matsuoka, Satoshi. Word Embeddings, Analogies, and Machine Learning: Beyond king - man + woman = queen. Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers. 2016

  234. [243]

    Proceedings of the 31st International Conference on Algorithmic Learning Theory , pages =

    What relations are reliably embeddable in Euclidean space? , author =. Proceedings of the 31st International Conference on Algorithmic Learning Theory , pages =. 2020 , editor =

  235. [244]

    A Primer in BERTology: What We Know About How BERT Works

    Anna Rogers and Olga Kovaleva and Anna Rumshisky. A Primer in BERTology: What We Know About How BERT Works. Transactions of the Association for Computational Linguistics. 2020. doi:10.1162/tacl_a_00349

Pith tools

Reviewed July 30, 2026 · model on record in the stance chip above.