{"id":"a6dfdd2b-5f4e-40db-a726-2c5849d65b7c","arxiv_id":"2607.26762","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Asymmetric semantic relations form clearer linear regions in LM spaces than symmetric ones, with only moderate encoding of directionality and transitivity and model-dependent reliance on lexical vs contextual cues.","lead":"Language-model embedding spaces encode asymmetric word relations (like hypernymy) more clearly than symmetric ones (like synonymy), and only moderately capture properties such as directionality and transitivity. The finding challenges the idea that distributional learning alone yields full semantic-relation knowledge and shows model type changes whether surface form or context matters more.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Modest linear-probe scores plus unmatched models undercut the leap from relative asymmetric/symmetric gaps to a distributional-semantics counter-claim.","rationale":"The reader correctly isolates the probe-as-geometry assumption (Method §3 + Limitations §7) as the load-bearing soft spot and already grades the paper CONDITIONAL for precisely the modest absolute effects, model confounds, and over-reaching distributional rhetoric. My stress test finds the same hinge: the anti-leakage design and multi-relation scope are genuine strengths, the asymmetric-vs-symmetric and distance-decay patterns are internally consistent, yet they remain relative differences among weak linear signals. No independent positive control or non-linear check is supplied, so the interpretive leap stays under-supported. That does not overturn the empirical package or require a harsher verdict; it simply confirms the reader’s CONDITIONAL call. Hence UNCHANGED and full agreement.","tokens_in":25201,"tokens_out":578,"duration_ms":28391,"concrete_test":"Train an otherwise identical non-linear (2-layer MLP) multi-relation probe on the same lemma-held-out folds and random controls; if controlled F_r for SYN/ANT rises above 0.40 while the asymmetric–symmetric gap collapses, or if matched-scale CLM/MLM pairs erase LLaMA’s advantage, the linear-geometry + distributional-counterclaim interpretation weakens materially.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (asymmetric relata regions clearer than symmetric; properties only moderately encoded; hence relations not uniformly learnable from distribution alone) rests on controlled bilinear-probe F/D/T scores (Eqs. 2–9, Tables 4–6, Fig. 2) being a faithful readout of linear relata-region geometry in final-layer space. Absolute controlled F_r never exceeds ~0.27 (asymmetric) / ~0.16 (symmetric), D_r ≤ 0.18, and long-distance T falls to chance or below; random-vector controls and lemma-held-out folds remove some confounds, yet leave open that (a) the geometry is simply non-linear (Limitations §7), (b) final-layer representations are the wrong locus, or (c) scale/data/objective confounds among ModernBERT / LLaDA / LLaMA / fastText (Table 3) drive the gaps rather than distributional learnability per se. Without a positive control that recovers high scores when linear geometry is known to be present, the relative pattern cannot securely license the strong theoretical suggestion in the abstract and §5.5–6.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper asks whether six WordNet semantic relations are encoded as linearly separable “relata regions” in LM token spaces, and whether those regions reflect symmetry/asymmetry and hypernymy transitivity. It trains bilinear multi-relation probes (Eqs. 2–4) on final-layer representations from ModernBERT, LLaDA-8B, LLaMA-3.1-8B, and a fastText baseline, using SemCor/WordNet triplets with lemma-held-out folds, removal of property-exhibiting pairs from training, an unrelated control class, random-representation controls, and multi-trial CV with Bonferroni-corrected tests. It further ablates lexical vs. contextual information via masking/previous-token states and degenerate/egalitarian attention. Controlled scores show clearer (still modest) separability for asymmetric than symmetric relations, only moderate directionality and short-distance transitivity, and model-dependent reliance on lexical vs. contextual cues. The authors read this as evidence that relation geometry is uneven and that semantic relations are not uniformly learnable from distributional information alone.","tokens_in":25583,"tokens_out":1597,"duration_ms":43919,"significance":"The work usefully broadens relation probing beyond hypernymy, introduces property-aware metrics (directionality; leakage-blocked transitivity), a geometrically interpretable bilinear probe, and a quantitative lexical/contextual ablation (ΔLEX/ΔCTX). Methodological care—lemma anti-leakage splits, property-pair removal, random probes, multi-trial CV, and transparent reporting of low absolute controlled scores—is a genuine strength and raises the bar relative to much of the prior probing literature. If the relative asymmetric/symmetric pattern and the information-source differences hold under tighter controls, the paper supplies concrete empirical constraints for distributional-semantics theory and for how different LM objectives encode relational structure. The contribution is primarily empirical and methodological rather than a decisive theoretical refutation.","major_comments":[{"comment":"The central theoretical suggestion (abstract; §5.5–6)—that uneven relation geometry shows semantic relations are not uniformly learnable from distribution alone—overreaches the absolute evidence. Controlled F_r never exceeds ~0.27 (asymmetric) / ~0.16 (symmetric) (Table 4); D_r ≤ 0.18 (Table 5); long-distance T falls to chance or below (Fig. 2, Table 6). Relative gaps are real under the authors’ controls, but without a positive control showing that the same bilinear probe recovers high controlled scores when linear relata-region geometry is independently known to be present, modest absolute performance cannot securely license a counter-claim against distributional semantics. Temper the claim to what the relative pattern supports, or add such a control.","section":"Abstract; §5.1–5.3; §5.5–6; Tables 4–6; Fig. 2"},{"comment":"Model comparison confounds scale, data, and objective (Table 3: ModernBERT 395M vs LLaMA/LLaDA 8B; unmatched corpora/context windows). LLaMA’s advantage on asymmetric F/D/T is repeatedly attributed in part to the causal objective and bidirectional vs unidirectional context (§5.5), yet Limitations §7 correctly notes these factors are entangled. As written, the abstract and discussion still invite a family-level conclusion (CLM vs MLM/DLM; lexical vs contextual importance). Either match models more carefully, add ablations that isolate objective/scale, or systematically downgrade causal language about “causal vs masked/diffusion” throughout results and conclusion.","section":"Table 3; §5.1–5.4; §5.5; §6–7"},{"comment":"The operational definition of “relation geometry” is linear separability under the bilinear probe (Eqs. 2–5, §3.1). Limitations §7 acknowledges that non-linear geometry is untested and that probing ≠ use. Because the negative findings on synonymy/antonymy and weak long-distance transitivity are load-bearing for the “not equally well-represented / not uniformly learnable” claim, the paper should either (i) include at least one non-linear probe baseline as a sensitivity check, or (ii) consistently frame all conclusions as about linear relata regions in final-layer space rather than about relation geometry or distributional learnability in general. The current framing oscillates between these.","section":"§3.1 Eqs. 2–5; §5.5–6; §7"},{"comment":"Transitivity evaluation trains only on direct pairs and tests indirect pairs (good anti-leakage design), but performance collapse with distance (Fig. 2) is also consistent with sense/representation drift and decreasing semantic overlap, not only with absence of transitive geometry. The discussion (§5.5) notes contextual dissimilarity for long-distance pairs; that alternative should be quantified (e.g., baseline similarity or probe confusion as a function of path length) so that “transitivity not encoded” is distinguished from “indirect pairs are simply harder under the same linear readout.”","section":"§3.2.3; §4.4; Fig. 2; §5.3; §5.5"}],"minor_comments":[{"comment":"Placeholder metadata remains in the front matter (“Action editor: {action editor name}”; “Submission received: DD Month YYYY”). Clean before any revision cycle.","section":"Front matter"},{"comment":"Running headers still say “Author’s Surnames Here / Running Article Title Here”.","section":"Headers"},{"comment":"Table 1 token counts and the 6.8M indirect-triplet figure in the intro should be cross-checked for consistency with the filtering narrative (intra-sentential removal, distance cutoffs).","section":"§4.2; Table 1; Table 2"},{"comment":"Figure 1 is schematic only; a 2D illustration of real probe hyperplanes or a qualitative PCA/UMAP of one target’s relata would help readers judge what “region” means empirically.","section":"Figure 1; §3.1"},{"comment":"Appendix C layer analysis is valuable but under-discussed in the main text; a short pointer in §5 on whether upper-layer dominance holds for all metrics would strengthen the final-layer choice.","section":"§4.5; Appendix C"},{"comment":"Minor wording: “summay relation predictabilities” (§3.2.1); “egalitarian decontextualisation” notation could be introduced once with a short pseudocode block for reproducibility.","section":"§3.2.1; §3.4"}],"recommendation":"major_revision","confidential_remarks":"The empirical package is careful and publishable after the interpretive overreach and model-confound issues are fixed; I would not reject. Fit for a CL journal is good (probing + lexical semantics). No integrity concerns. The skeptic note’s worry about linear-probe faithfulness and unmatched models is substantive and should be addressed in revision; it does not, in my view, sink the relative-pattern contribution."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is a careful comparative measurement: six WordNet relations plus UNR, bilinear multi-class probes, directionality under order-swap, and distance-stratified transitivity trained only on direct pairs. That package is cleaner than most prior hypernymy probes.\n\nWhat is actually new is the breadth plus the hygiene. Lemma-held-out folds, removal of property-exhibiting pairs from training, random-representation controls, multi-trial CV with Bonferroni-corrected permutation tests, and the ΔLEX/ΔCTX ablations are done properly. They report low absolute controlled scores honestly (F_r mostly under 0.27, D_r ≤ 0.18, long-distance T to chance). The relative pattern—asymmetric relations clearer than synonymy/antonymy, short-distance transitivity better than long—is consistent across ModernBERT, LLaDA, and LLaMA. The lexical-vs-contextual split by model family is a genuine secondary result.\n\nSoft spots in proportion: absolute effect sizes stay small, so the leap to “counter-evidence that relations are not uniformly learnable from distribution alone” is stronger than the linear final-layer signal supports. They flag the linear-probe limit themselves. The three LMs are unmatched on scale and data, so LLaMA’s edge is confounded; the static baseline winning on antonymy is interesting but under-interpreted. No positive control where known linear geometry should recover high scores. None of this sinks the empirical core.\n\nMath and citation pattern look fine—bilinear scoring is standard, WordNet/SemCor grounding is external, related work covers the right hypernymy and synonymy papers. No circularity problem.\n\nThis is for people who care about probing, lexical semantics, and what distributional models actually encode. Worth a serious referee. I would engage: cite the design and the relative pattern, push back on the theoretical framing in revision.","headline":"Solid multi-relation geometry audit with unusually clean anti-leakage design; the asymmetric-vs-symmetric pattern is real, but absolute scores are modest and the distributional-hypothesis punchline outruns the linear-probe evidence.","tokens_in":26155,"tokens_out":505,"would_cite":true,"duration_ms":17591,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Semantic relations do not leave equally clear footprints in language-model vector space; asymmetric ones are clearer than symmetric ones.","keywords":["semantic relations","relation geometry","language models","probing","distributional semantics","hypernymy","synonymy","lexical vs contextual information"],"falsifier":"Under the same lemma-held-out, leakage-blocked setup, synonym or antonym pairs becoming as linearly separable and directionally consistent as hypernym or hyponym pairs, or long-distance hypernym chains remaining highly predictable when the probe is trained only on direct pairs.","tokens_in":26089,"feed_emoji":"📐","tokens_out":869,"duration_ms":25857,"temperature":0.7,"pith_summary":"Language models turn words into vectors, but it is unclear how much of the structure of semantic relations—hypernymy, synonymy, antonymy, and the rest—is actually present as geometry in those spaces. This paper tests that question directly: whether words related to a target by the same relation cluster in a shared region, whether those regions encode classic properties such as asymmetry and transitivity, and whether the signal comes more from word form or from context. Across causal, masked, and diffusion models, asymmetric relations show relatively clearer separable regions than symmetric ones, while directionality and long-distance transitivity are only moderately recovered. The pattern is uneven enough to suggest that not every semantic relation is equally learnable from distributional information alone. A sympathetic reader cares because this is a concrete check on a foundational claim of distributional semantics, not just another leaderboard number.","feed_headline":"LM vector spaces encode relations unevenly","feed_subtitle":"Asymmetric links leave clearer regions than synonyms; long-distance hierarchy stays weak.","key_machinery":"A bilinear multi-relation probe that scores whether a target–relatum pair falls into a relation-specific linear “relata region,” evaluated with controlled predictability, directionality, and transitivity metrics, plus ablations that strip lexical form or correct context.","core_discovery":"Relation geometry in language-model semantic space is not uniform: relata of asymmetric relations occupy relatively distinct linear regions, while symmetric relations do not, and properties such as directionality and especially long-distance transitivity are only moderately encoded—evidence that semantic relations are not equally recoverable from distributional learning alone.","pith_inferences":["If linear relata regions are weak for similarity but stronger for relatedness, many “analogy” or vector-offset recipes may be succeeding on relatedness structure rather than true synonymy.","The gap between short- and long-distance transitivity suggests hierarchical resources still need explicit structure beyond what next-token or masked training induces.","A natural next test is whether non-linear probes close the synonymy gap or merely confirm that the missing structure is absent, not just hard to read linearly.","Training objectives that mix shared-neighbour and co-occurrence signals may explain why LMs beat static embeddings on asymmetric relations but not on antonymy."],"forward_implications":["Asymmetric relatedness (hypernymy, meronymy and their reverses) should be easier to read out of LM embeddings than synonymy-style similarity.","Probes and applications that assume uniform relational geometry across WordNet-style relations will systematically overestimate what distributional spaces encode for synonyms and antonyms.","Causal models will lean more on surface form for relational geometry, while masked and diffusion models will lean more on context—except for indirect hypernymy, where lexical form matters more across model types.","Long-distance hierarchical inference cannot be treated as a free consequence of local hypernym geometry learned from co-occurrence alone."],"fun_headline_variants":["LM spaces encode asymmetric relations more clearly than symmetric ones","Relata regions form unevenly across relation types in LM geometry","Asymmetric links carve distinct regions; symmetry stays blurry","Relation geometry in LMs favors direction over long-range hierarchy","Distributional spaces recover some relations better than others"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That how well this linear probe separates held-out word pairs is a faithful readout of whether relation-specific regions and relation properties actually exist in the model’s space.","fun_headline_variants_meta":{"raw":{"variants":["LM spaces encode asymmetric relations more clearly than symmetric ones","Relata regions form unevenly across relation types in LM geometry","Asymmetric links carve distinct regions; symmetry stays blurry","Relation geometry in LMs favors direction over long-range hierarchy","Distributional spaces recover some relations better than others"]},"model":"grok-4.5","effort":"low","cost_usd":0.003806,"raw_usage":{"total_tokens":1210,"prompt_tokens":809,"num_sources_used":0,"completion_tokens":63,"cost_in_usd_ticks":38064000,"prompt_tokens_details":{"text_tokens":809,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":338,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":809,"tokens_out":63,"duration_ms":7128,"temperature":1.0,"reasoning_tokens":338,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T21:50:26.729692+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Under the same lemma-held-out, leakage-blocked setup, synonym or antonym pairs becoming as linearly separable and directionally consistent as hypernym or hyponym pairs, or long-distance hypernym chains remaining highly predictable when the probe is trained only on direct pairs.","supporting_citations":[],"review_version":1}