Pith. sign in

REVIEW 2 major objections 1 minor 18 references

Geometry of Semantic Space: Comparative Study of Discrete and Continuous Models

T0 review · 2 major / 1 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read Lexical co-occurrence graphs and transformer embeddings produce similar local but different overall topologies on French debate texts.

desk verdict The paper's claim of similar local but different global topology between CamemBERT and co-occurrence graphs on one French corpus rests on an unexamined premise with no quantitative backing. read the letter →

arxiv 2606.07183 v1 pith:LQFSD3SJ submitted 2026-06-05 cs.CL

classification cs.CL
keywords semanticgeometrylexicalco-occurrencegraphstransformerembeddingstopologycomparisonFrenchdebatecorpusspacegraph-basedmodelsNLPmodel
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to compare the geometries that arise when semantic relations are modeled by supervised vector embeddings such as CamemBERT versus when they are modeled by lexical co-occurrence graphs. On the French Great National Debate corpus the two representations display matching local topology yet markedly different global structure, with the graphs producing a clearer, more readable organization of meaning. A reader would care because the contrast points to concrete ways the strengths of each approach might be combined. The authors therefore propose using graph structure as a reference that can steer neural models toward more stable and interpretable convergence.

What carries the argument

Methodology for side-by-side analysis of graph structure versus embedding topology applied to the same corpus.

What would settle it

Repeating the comparison on the same corpus with a different graph-construction rule or a differently trained embedding model and still obtaining matching overall topologies would falsify the claim of very different global structures.

Watch

Extended reading notes

Core claim

When the structure of lexical co-occurrence graphs is compared with the topology of embeddings induced by supervised transformers on the French Great National Debate corpus, local topology is similar but overall structure and topology differ substantially. Graph-based models exhibit clearer and more human-readable organization of meaning, whereas transformer embeddings display unsatisfactory distributions. These observations indicate that the two modeling families supply complementary perspectives and open a route to guide neural architectures toward convergence with graph structures.

Load-bearing premise

Lexical co-occurrence graphs encode semantic relations more directly than the embeddings produced by supervised transformers.

Editorial extensions

If this is right

  • Deep supervised models and graph-based models supply complementary views of semantic organization.
  • Graph structures can serve as a reference that guides neural architectures toward more stable and interpretable forms.
  • Local similarity combined with global divergence implies that the two approaches capture semantic relations at different scales.
  • Clearer organization in graphs offers a concrete benchmark for improving the distributional properties of transformer embeddings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Testing the same comparison on corpora from other languages or domains would show whether the local-global split is specific to French debate texts.
  • Directly incorporating co-occurrence graph constraints into transformer training loops could be examined as one route to the convergence the paper envisions.
  • The global-topology difference may help explain why transformer representations sometimes exhibit less stability across fine-tuning runs.
  • Mapping particular semantic relations onto the local versus global features of each geometry could refine how each model is used for downstream tasks.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper compares the semantic geometries of supervised transformer embeddings (e.g., CamemBERT) and lexical co-occurrence graphs on the French 'Great National Debate' corpus. It asserts that co-occurrence graphs encode semantic relations more directly, that graph-based models yield clearer human-readable organization, and that the two approaches exhibit similar local topology but markedly different global structure and topology, thereby offering complementary perspectives for improving neural architectures.

Significance. If the comparison were placed on a sound methodological footing with explicit controls and quantitative metrics, the work could usefully illustrate how discrete graph structures might stabilize or interpret continuous embedding spaces. At present the absence of any reported measures prevents assessment of whether the claimed complementarity is robust or merely an artifact of mismatched representation formats.

major comments (2)
  1. [Abstract] Abstract: The premise that 'lexical co-occurrence graphs that encode semantic relations more directly' than CamemBERT embeddings is stated without any definition of 'directly,' any alignment procedure between discrete graphs and continuous high-dimensional vectors, or any control for supervision and training-objective differences. This premise is load-bearing for the claim that observed global-topology differences can be attributed to model class rather than to inductive-bias mismatch.
  2. [Abstract] Abstract: The central empirical result ('similar local topology but a very different overall structure and topology') is asserted without quantitative measures, error bars, dataset statistics, exclusion criteria, or even a description of the topology metrics employed, rendering the claim unevaluable.
minor comments (1)
  1. [Abstract] Abstract, final sentence: 'Theses findings' is a typographical error for 'These findings.'

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive comments, which identify key areas where the manuscript can be strengthened methodologically. We respond to each major comment below and will revise the paper accordingly.

read point-by-point responses
  1. Referee: [Abstract] Abstract: The premise that 'lexical co-occurrence graphs that encode semantic relations more directly' than CamemBERT embeddings is stated without any definition of 'directly,' any alignment procedure between discrete graphs and continuous high-dimensional vectors, or any control for supervision and training-objective differences. This premise is load-bearing for the claim that observed global-topology differences can be attributed to model class rather than to inductive-bias mismatch.

    Authors: We agree that the phrasing requires clarification to avoid ambiguity. In revision we will define 'more directly' explicitly as construction via raw co-occurrence counts without parametric learning. We will also add a description of the alignment procedure (mapping both representations into a shared topological space via graph embedding and dimensionality reduction) and a new subsection discussing supervision and objective differences as potential confounds. These additions will allow readers to assess whether topology differences are attributable to model class. revision: yes

  2. Referee: [Abstract] Abstract: The central empirical result ('similar local topology but a very different overall structure and topology') is asserted without quantitative measures, error bars, dataset statistics, exclusion criteria, or even a description of the topology metrics employed, rendering the claim unevaluable.

    Authors: We accept that the current presentation lacks sufficient quantitative detail for evaluation. Although the manuscript describes the metrics employed, the revision will incorporate explicit numerical results (including error bars), full dataset statistics for the Great National Debate corpus, node/edge exclusion criteria, and a precise account of all local and global topology metrics. This will render the central claim fully evaluable. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical comparison lacks any derivation chain or fitted inputs

full rationale

The paper conducts a direct empirical comparison of local/global topology between CamemBERT embeddings and lexical co-occurrence graphs on the Great National Debate corpus. No equations, parameter fitting, predictions derived from fits, or self-citation chains appear in the provided text. The premise that graphs encode semantics 'more directly' is an unexamined modeling choice but does not reduce any claimed result to its own inputs by construction. The reported similarity in local topology and difference in overall structure is presented as an observed outcome of the analysis, not a tautology. This is a standard self-contained empirical study.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Abstract supplies no information on free parameters, background axioms, or new postulated entities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Geometry of Semantic Space: Comparative Study of Discrete and Continuous Models." pith.science (2026). https://pith.science/paper/LQFSD3SJ

@misc{pith2026260607183,
  author       = {Pith},
  title        = {Pith review of: Geometry of Semantic Space: Comparative Study of Discrete and Continuous Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LQFSD3SJ}},
  note         = {Machine review of arXiv:2606.07183}
}
read the original abstract

This work examines the semantic geometry underlying NLP models. We compare supervised vector embeddings, such as CamemBERT, with lexical co-occurrence graphs that encode semantic relations more directly. While transformer-based embeddings achieve strong performance, their induced geometries often display unsatisfactory distributions. In contrast, graph-based models reveal a clearer and more human-readable organization of meaning. We have implemented a methodology that allows us to perform a comparative analysis either based on the structure of the graphs or based on the topology of the embeddings induced by these two approaches. The results of the comparison -- applied to the French "Great National Debate" corpus a collection of citizen contributions to the public debate -- show a similar local topology but a very different overall structure and topology. Theses findings suggest complementary perspectives between deep supervised models and graph-based models, considering a new pathway to guide neural architectures toward more stable and interpretable convergence with graphs structures.

Figures

Figures reproduced from arXiv: 2606.07183 by the authors.

Figure 2
Figure 2. Correlation between metric distances in the [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Distribution of cosine similarities between vectors of the four types of embeddings of [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Distribution of cosine similarities between pairs of embeddings of cliques from the CamemBERT model. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figures from the paper (3 more)
Figure 5
Figure 5. Figure 5: Gc (blue) and Gb (orange) degree (top) and betweenness centrality (bottom) distributions. Gb, with the clusters colored. Overlaid on each partition is a histogram of cluster sizes. We observe strong heterogeneity in Gb : few large clusters and many small ones. The hist…
Figure 6
Figure 6. Figure 6: Upper, Gb (left) and Gc (right) colored according to the Infomap partition of each graph. Graphs are constructed at thPMI = 9, p = 0.05% and sv = seucl, FD-type embedding. Lower, the respective distribution of cluster sizes in the associated Infomap partition. Cluster …
Figure 7
Figure 7. Figure 7: Evolution of trustworthiness at k = 10 on the Infomap partitioning as a function of the dimension of the embedding of Gc at p = 0.1% across the four types of embeddings. were to prove true, they would result not only in quantitative and qualitative gains but also in an…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

18 extracted references · 3 canonical work pages

  1. [1]

    Aho and Jeffrey D

    Alfred V. Aho and Jeffrey D. Ullman , title =. 1972

  2. [2]

    Zenodo , year=

    Simplemma , author=. Zenodo , year=

  3. [3]

    Publications Manual , year = "1983", publisher =

  4. [4]

    Proceedings of Machine Translation Summit IX: Papers , year=

    Lexical knowledge representation with contextonyms , author=. Proceedings of Machine Translation Summit IX: Papers , year=

  5. [5]

    Bert: Pre-training of deep bidirectional transformers for language understanding , author=. Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) , pages=

  6. [6]

    RoBERTa: A Robustly Optimized BERT Pretraining Approach

    Roberta: A robustly optimized bert pretraining approach , author=. arXiv preprint arXiv:1907.11692 , year=

  7. [7]

    Proceedings of the 58th annual meeting of the association for computational linguistics , pages=

    CamemBERT: a tasty French language model , author=. Proceedings of the 58th annual meeting of the association for computational linguistics , pages=

  8. [8]

    Sentence-bert: Sentence embeddings using siamese bert-networks , author=. Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP) , pages=

Show all 18 references
  1. [9]

    Chandra and Dexter C

    Ashok K. Chandra and Dexter C. Kozen and Larry J. Stockmeyer , year = "1981", title =. doi:10.1145/322234.322243

  2. [10]

    arXiv preprint arXiv:2505.22107 , year=

    Curse of High Dimensionality Issue in Transformer for Long-context Modeling , author=. arXiv preprint arXiv:2505.22107 , year=

  3. [11]

    PloS one , volume=

    Multilevel compression of random walks on networks reveals hierarchical organization in large integrated systems , author=. PloS one , volume=. 2011 , publisher=

  4. [12]

    Proceedings of the 22nd ACM SIGKDD international conference on Knowledge discovery and data mining , pages=

    node2vec: Scalable feature learning for networks , author=. Proceedings of the 22nd ACM SIGKDD international conference on Knowledge discovery and data mining , pages=

  5. [13]

    science , volume=

    A global geometric framework for nonlinear dimensionality reduction , author=. science , volume=. 2000 , publisher=

  6. [14]

    Congressus numerantium , volume=

    A heuristic for graph drawing , author=. Congressus numerantium , volume=

  7. [15]

    Scalable training of

    Andrew, Galen and Gao, Jianfeng , booktitle=. Scalable training of

  8. [16]

    Dan Gusfield , title =. 1997

  9. [17]

    Tetreault , title =

    Mohammad Sadegh Rasooli and Joel R. Tetreault , title =. Computing Research Repository , volume =. 2015 , url =

  10. [18]

    A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =

    Ando, Rie Kubota and Zhang, Tong , Issn =. A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =. Journal of Machine Learning Research , Month = dec, Numpages =

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.