Pith. sign in

REVIEW 5 minor 56 references

Towards Non-Euclidean Foundation Models: Advancing AI Beyond Euclidean Frameworks

T0 review · 0 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This workshop proposal argues that embedding foundation models in curved, non-Euclidean spaces could improve search, recommendation, and content understanding on the web.

desk verdict A well-organized workshop proposal on non-Euclidean foundation models, not a research paper—plausible motivation, no testable claims, so unverdictable as a scientific contribution. read the letter →

arxiv 2505.14417 v1 pith:R3GEGYYU submitted 2025-05-20 cs.CG cs.LG

classification cs.CGcs.LG
keywords FoundationModelsNon-EuclideanGeometryHyperbolicSpaceInformationRetrievalGeometricLearningRepresentationRecommenderSystemsWeb-scaleApplications
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a workshop proposal arguing that the near-universal use of Euclidean geometry in foundation models is a fundamental limitation, and that non-Euclidean spaces—hyperbolic, spherical, and mixed-curvature—offer a better fit for web data such as social networks, query-document pairs, and user-item interactions. The authors are trying to establish a research agenda: if foundation models are rethought in curved geometries, search, recommendation, and content understanding should improve. They do not run experiments; the claim rests on prior non-Euclidean representation results at small scale plus a call for new architectures, benchmarks, and robustness studies. A sympathetic reader cares because every modern large language model lives in Euclidean embedding space, while the structures it must model are often hierarchical and graph-like.

What carries the argument

The central machinery is the curvature of the representation space. Hyperbolic space, the negatively curved model in which tree-like hierarchies embed with exponentially growing volume, is proposed for taxonomies and social networks; spherical (positively curved) space is proposed for cyclic and symmetric structures; mixed-curvature product spaces combine both to match heterogeneous graphs. The paper treats these spaces as the geometric inductive bias that carries the benefit: matching the space's curvature to the data's intrinsic structure is what should make embeddings more efficient and more accurate.

What would settle it

Train the same transformer architecture on the same web-scale corpus in Euclidean and hyperbolic geometry with matched compute, then compare query-document retrieval and recommendation accuracy; if the hyperbolic model does not beat the Euclidean baseline, the paper's central premise is refuted.

Watch

Extended reading notes

Core claim

The paper's central claim is that Euclidean space is not the right geometric default for foundation models that work with web data. It asserts that hyperbolic, spherical, and mixed-curvature spaces can represent hierarchical semantics, network topology, query-document similarity, and user-item relationships more faithfully, and that integrating these geometries into foundation models would yield measurable gains in search, recommendation, and content understanding. The paper advances this claim as a program for the community: it defines the scope of a workshop, invites submissions on theory, architectures, applications, trustworthiness, and benchmarks, and positions the workshop as the first web-centric event on non-Euclidean foundation models.

Load-bearing premise

The paper assumes that the benefits shown for non-Euclidean embeddings on small graph and representation-learning benchmarks will survive at the scale of foundation-model training on web-scale data, without testing that transfer anywhere.

Editorial extensions

If this is right

  • According to the paper, language models that embed text in hyperbolic space should represent hierarchies with fewer dimensions, yielding more compact and more accurate semantic search.
  • The paper expects recommender systems to treat user-item interaction graphs as curved, capturing higher-order relationships without the distortion Euclidean embeddings incur.
  • Mixed-curvature models, the paper argues, could handle web data that mixes tree-like and cyclic structure, such as knowledge graphs with both taxonomic and relational edges.
  • The paper calls for new web-scale benchmarks and evaluation protocols, since current non-Euclidean evidence comes from small-scale tasks.
  • Robustness and trustworthiness studies are needed before curved models are deployed, because adversarial and privacy behavior in non-Euclidean space is largely unexplored.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the same logic would apply to multimodal web content, where aligning image, text, and graph modalities in a common curved space could reduce the distortion that flat alignment introduces.
  • Editorial inference: a cheap test of the scale-transfer premise is to take an existing non-Euclidean recommender and evaluate it on a web-scale interaction graph; the paper does not include such a test.
  • Editorial inference: if curved spaces become standard, the field will need numerical-stability and optimization tools tailored to non-Euclidean manifolds at foundation-model scale, an area the paper lists but does not develop.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. This manuscript is a workshop summary/proposal for a full-day workshop at WWW 2025 on Non-Euclidean Foundation Models and Geometric Learning (NEGEL). It motivates the workshop by arguing that Euclidean embeddings are a limiting default for foundation models and that non-Euclidean (hyperbolic, spherical, mixed-curvature) representations have shown benefits on web-related tasks such as search, recommendation, and content understanding. The document provides a detailed schedule, five topic areas, organizer and speaker biographies, a program committee list, and a diversity statement. No experiments, derivations, or datasets are presented; the scientific claims are supported only by citations to prior work.

Significance. As a workshop proposal, the manuscript's significance lies in identifying an emerging and plausible research direction and in offering a concrete organizational structure for it. The organizers are well-known in the relevant fields, the invited speakers are appropriate, and the stated topics cover theory, algorithms, applications, trustworthiness, and benchmarks. If non-Euclidean foundation models deliver on their promise, this workshop could help coalesce a new subfield. However, no evidence is offered beyond citations, and the central premise—that benefits observed in small-scale non-Euclidean embeddings will persist at foundation-model scale—remains an open (and explicitly acknowledged) research question.

minor comments (5)
  1. [Abstract] 'mach ine learning' contains a stray space; should read 'machine learning'.
  2. [References] Reference [10] uses 'Vıctor' with a dotless 'ı'; this should be 'Víctor' (or the appropriate Unicode encoding) for correct typesetting.
  3. [Section 3.2] 'Smita Krishnaswamy is an Associate professor in Genetics' has inconsistent capitalization; 'Associate Professor' is the standard form.
  4. [Section 2] '20 ∼ 100 attendees' uses a tilde; consider an en-dash or the phrase '20 to 100' for consistency with the rest of the text.
  5. [Section 1 (Proposed Duration)] The scheduled time '8:00–9:00AM' is listed as 'Poster setup'; consider clarifying whether this is for participants as well, since the session list otherwise begins at 9:00AM.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity identified: the document is a workshop proposal with no derivation chain, fitted parameters, or testable predictions to reduce.

full rationale

The manuscript is a WWW Companion workshop proposal, not a research paper. Its central claim—that integrating non-Euclidean geometries with foundation models can improve web applications—is presented in Section 1 as a motivation, phrased as 'has great potential,' and is not derived from any equations, experiments, or fitted parameters. The remainder of the document describes scope, invited speakers, schedule, committee, and logistics. There is no derivation chain in which a prediction is equivalent to an input by construction, no parameter fitted to data and then renamed as a prediction, and no load-bearing self-citation invoked to force a conclusion. The cited prior works are used as supporting literature for the workshop's topic areas, not as inputs to a derived result. Consequently, none of the enumerated circularity patterns apply, and the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

All substantive claims are borrowed from cited literature; the paper's own contribution is organizational. The ledger lists the three key assumptions that the proposal depends on.

assumptions (3)
  • domain assumption Euclidean space is fundamentally limited for modeling complex relational and hierarchical data.
    Invoked in the abstract and Section 1; supported only by citations to [2,12,15,29].
  • domain assumption Non-Euclidean spaces provide more efficient and effective representations for web-related data.
    Invoked with citations [4,8,14,19,20,24,25,27,34,38,39,44,46,49]. This is a working hypothesis in the literature, not a theorem.
  • ad hoc to paper Integrating foundation models with non-Euclidean geometries will enhance their performance.
    This is the workshop's core thesis, stated in the abstract and Section 1. It is not proven in this document and is the main premise the workshop is built on.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Non-Euclidean Foundation Models: Advancing AI Beyond Euclidean Frameworks." pith.science (2026). https://pith.science/paper/R3GEGYYU

@misc{pith2026250514417,
  author       = {Pith},
  title        = {Pith review of: Towards Non-Euclidean Foundation Models: Advancing AI Beyond Euclidean Frameworks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/R3GEGYYU}},
  note         = {Machine review of arXiv:2505.14417}
}
read the original abstract

In the era of foundation models and Large Language Models (LLMs), Euclidean space is the de facto geometric setting of our machine learning architectures. However, recent literature has demonstrated that this choice comes with fundamental limitations. To that end, non-Euclidean learning is quickly gaining traction, particularly in web-related applications where complex relationships and structures are prevalent. Non-Euclidean spaces, such as hyperbolic, spherical, and mixed-curvature spaces, have been shown to provide more efficient and effective representations for data with intrinsic geometric properties, including web-related data like social network topology, query-document relationships, and user-item interactions. Integrating foundation models with non-Euclidean geometries has great potential to enhance their ability to capture and model the underlying structures, leading to better performance in search, recommendations, and content understanding. This workshop focuses on the intersection of Non-Euclidean Foundation Models and Geometric Learning (NEGEL), exploring its potential benefits, including the potential benefits for advancing web-related technologies, challenges, and future directions. Workshop page: [https://hyperboliclearning.github.io/events/www2025workshop](https://hyperboliclearning.github.io/events/www2025workshop)

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

56 extracted references · 50 canonical work pages

  1. [1]

    Nico Alvarado, Hans Lobel, and Mircea Petrache. 2023. Cu rvature-Dimension Tradeoff for Generalization in Hyperbolic Space. In NeurIPS 2023 Workshop

  2. [2]

    Gregor Bachmann, Gary Bécigneul, and Octavian Ganea. 20 20. Constant curva- ture graph convolutional networks. In ICML. PMLR, 486–496

  3. [3]

    Beatrice Bevilacqua, Fabrizio Frasca, Derek Lim, Balas ubramaniam Srinivasan, Chen Cai, Gopinath Balamurugan, Michael M Bronstein, and Ha ggai Maron

  4. [4]

    Luca Bombelli, Joohan Lee, David Meyer, and Rafael D Sork in. 1987. Space-time as a causal set. Physical review letters 59, 5 (1987), 521

  5. [5]

    Ines Chami, Zhitao Ying, Christopher Ré, and Jure Leskov ec. 2019. Hyperbolic graph convolutional neural networks. In NeurIPS. 4868–4879

  6. [6]

    Yankai Chen, Menglin Yang, Yingxue Zhang, Mengchen Zhao , Ziqiao Meng, Jianye Hao, and Irwin King. 2022. Modeling scale-free graph s with hyperbolic geometry for knowledge-aware recommendation. In WSDM. 94–102

  7. [7]

    Octavian Ganea, Gary Bécigneul, and Thomas Hofmann. 201 8. Hyperbolic en- tailment cones for learning hierarchical embeddings. In ICML. PMLR, 1646– 1655

  8. [8]

    Albert Gu, Frederic Sala, Beliz Gunel, and Christopher R é. 2019. Learning mixed- curvature representations in product spaces. In ICLR

Show all 56 references
  1. [9]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 20 16. Deep residual learning for image recognition. In CVPR. 770–778

  2. [10]

    Emiel Hoogeboom, Vıctor Garcia Satorras, Clément Vign ac, and Max Welling

  3. [11]

    Valentin Khrulkov, Leyla Mirvakhabova, Evgeniya Usti nova, Ivan Oseledets, and Victor Lempitsky. 2020. Hyperbolic image embeddings. In CVPR. 6418–6428

  4. [12]

    Bobak T Kiani, Thien Le, Hannah Lawrence, Stefanie Jege lka, and Melanie We- ber. 2024. On the hardness of learning under symmetries. ICLR (2024)

  5. [13]

    Bobak T Kiani, Jason Wang, and Melanie Weber. 2024. Hard ness of Learning Neural Networks under the Manifold Hypothesis. In NeurIPS

  6. [14]

    Marc Law and Jos Stam. 2020. Ultrahyperbolic represent ation learning. In NeurIPS. 1668–1678

  7. [15]

    Nathan Linial, Eran London, and Yuri Rabinovich. 1995. The geometry of graphs and some of its algorithmic applications. Combinatorica 15 (1995), 215–245

  8. [16]

    Jiahong Liu, Xinyu Fu, Menglin Yang, Weixi Zhang, Rex Yi ng, and Irwin King

  9. [17]

    Jiahong Liu, Menglin Yang, Min Zhou, Shanshan Feng, and Philippe Fournier- Viger. 2022. Enhancing hyperbolic graph embeddings via con trastive learning. arXiv:2201.08554 (2022)

  10. [18]

    Qi Liu, Maximilian Nickel, and Douwe Kiela. 2019. Hyper bolic graph neural networks. In NeurIPS. 8230–8241

  11. [19]

    Yu Meng, Jiaxin Huang, Guangyuan Wang, Chao Zhang, Hong lei Zhuang, Lance Kaplan, and Jiawei Han. 2019. Spherical text embedding. In NeurIPS

  12. [20]

    Pascal Mettes, Mina Ghadimi Atigh, Martin Keller-Ress el, Jeffrey Gu, and Ser- ena Yeung. 2023. Hyperbolic Deep Learning in Computer Visio n: A Survey. arXiv:2305.06611 (2023)

  13. [21]

    Gal Mishne, Zhengchao Wan, Yusu Wang, and Sheng Yang. 20 23. The numerical stability of hyperbolic representation learning. In ICML. PMLR, 24925–24949

  14. [22]

    Maximillian Nickel and Douwe Kiela. 2017. Poincaré emb eddings for learning hierarchical representations. In NeurIPS. 6338–6347

  15. [23]

    Maximillian Nickel and Douwe Kiela. 2018. Learning Con tinuous Hierarchies in the Lorentz Model of Hyperbolic Geometry. In ICML. 3779–3788

  16. [24]

    Wei Peng, Tuomas Varanka, Abdelrahman Mostafa, Hengli n Shi, and Guoying Zhao. 2021. Hyperbolic deep neural networks: A survey. TPAMI (2021)

  17. [25]

    Aaron Sim, Maciej L Wiatrak, Angus Brayne, Páidí Creed, and Saee Paliwal. 2021. Directed graph embeddings in pseudo-riemannian manifolds . In ICML. PMLR, 9681–9690

  18. [26]

    Li Sun, Zhenhao Huang, Suyang Zhou, Qiqi Wan, Hao Peng, a nd Philip Yu. 2025. RiemannGFM: Learning a Graph Foundation Model from Riemannian Geometry. arXiv:2502.03251 (2025)

  19. [27]

    Li Sun, Zhongbao Zhang, Junda Ye, Hao Peng, Jiawei Zhang , Sen Su, and Philip S Yu. 2022. A Self-supervised Mixed-curvature Graph Neural N etwork. AAAI (2022)

  20. [28]

    Atsushi Suzuki, Atsushi Nitanda, Taiji Suzuki, Jing Wa ng, Feng Tian, and Kenji Yamanishi. 2023. Tight and fast generalization error bound of graph embedding in metric space. arXiv:2305.07971 (2023)

  21. [29]

    Atsushi Suzuki, Atsushi Nitanda, Jing Wang, Linchuan X u, Kenji Yamanishi, and Marc Cavazza. 2021. Generalization Error Bound for Hyperbolic Ordinal Embed- ding. In ICML. PMLR, 10011–10021

  22. [30]

    Yi Tay, Luu Anh Tuan, and Siu Cheung Hui. 2018. Hyperboli c representation learning for fast and efficient neural question answering. In WSDM. 583–591

  23. [31]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288 (2023)

  24. [32]

    Max van Spengler, Erwin Berkhout, and Pascal Mettes. 20 23. Poincare resnet. In ICCV. 5419–5428

  25. [33]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. At tention is all you need. In NeurIPS. Curran Associates, Inc., 1–11

  26. [34]

    Shen Wang, Xiaokai Wei, Cicero Nogueira Nogueira dos Sa ntos, Zhiguo Wang, Ramesh Nallapati, Andrew Arnold, Bing Xiang, Philip S Yu, an d Isabel F Cruz

  27. [35]

    Yujie Wang, Shuo Zhang, Junda Ye, Hao Peng, and Li Sun. 20 24. A Mixed- Curvature Graph Diffusion Model. In CIKM. 2482–2492

  28. [36]

    Melanie Weber. 2020. Neighborhood Growth Determines G eometric Priors for Relational Representation Learning. In AISTATS, Vol. 108. 266–276

  29. [37]

    Melanie Weber and Suvrit Sra. 2023. Global optimality f or Euclidean CCCP under Riemannian convexity. In ICML

  30. [38]

    Mixed-curvature multi-relational graph neural netw ork for knowledge graph completion. In WWW. 1761–1771

  31. [39]

    Bo Xiong, Shichao Zhu, Nico Potyka, Shirui Pan, Chuan Zh ou, and Steffen Staab

  32. [40]

    Menglin Yang, Aosong Feng, Bo Xiong, Jihong Liu, Irwin K ing, and Rex Ying

  33. [41]

    Menglin Yang, Zhihao Li, Min Zhou, Jiahong Liu, and Irwi n King. 2022. Hicf: Hyperbolic informative collaborative filtering. In KDD. 2212–2221

  34. [42]

    Melanie Weber, Manzil Zaheer, Ankit Singh Rawat, Adity a K Menon, and San- jiv Kumar. 2020. Robust large-margin learning in hyperboli c space. In NeurIPS. 17863–17873

  35. [43]

    Menglin Yang, Min Zhou, Marcus Kalander, Zengfeng Huan g, and Irwin King

  36. [44]

    InNeurIPS

    Pseudo-riemannian graph convolutional networks. InNeurIPS. 3488–3501

  37. [45]

    Menglin Yang, Min Zhou, Jiahong Liu, Defu Lian, and Irwi n King. 2022. HRCF: Enhancing collaborative filtering via hyperbolic geometri c regularization. In WWW. 2462–2471

  38. [46]

    ArXiv abs/2410.04010 (2024)

    Hyperbolic Fine-tuning for Large Language Models. ArXiv abs/2410.04010 (2024)

  39. [47]

    Delvin Ce Zhang, Rex Ying, and Hady W Lauw. 2023. Hyperbo lic Graph Topic Modeling Network with Continuously Updated Topic Tree. In KDD. 3206–3216

  40. [48]

    Menglin Yang, Harshit Verma, Delvin Ce Zhang, Jiahong L iu, Irwin King, and Rex Ying. 2024. Hypformer: Exploring Efficient Transformer Fully in Hyperbolic Space. In KDD

  41. [49]

    Shichao Zhu, Shirui Pan, Chuan Zhou, Jia Wu, Yanan Cao, a nd Bin Wang. 2020. Graph geometry interaction learning. In NeurIPS. 7548–7558

  42. [50]

    Discrete-time temporal network embedding via implicit hierarchical learn- ing in hyperbolic space. In KDD. 1975–1985

  43. [51]

    Menglin Yang, Min Zhou, Zhihao Li, Jiahong Liu, Lujia Pa n, Hui Xiong, and Irwin King. 2022. Hyperbolic Graph Neural Networks: A Revie w of Methods and Applications. arXiv:2202.13852 (2022)

  44. [53]

    Dingyi Zhang, Yingming Li, and Zhongfei Zhang. 2020. De ep metric learning with spherical embedding. In NeurIPS. 18772–18783

  45. [55]

    Sixiao Zhang, Hongxu Chen, Xiao Ming, Lizhen Cui, Hongz hi Yin, and Guan- dong Xu. 2021. Where are we in embedding spaces?. In KDD. 2223–2231

  46. [2021]

    Equivariant Subgraph Aggregation Networks. In ICLR

  47. [2022]

    Equivariant diffusion for molecule generation in 3d. In ICML. PMLR, 8867– 8887

  48. [2024]

    In FedKDD workshop at KDD

    Client-Specific Hyperbolic Federated Learning. In FedKDD workshop at KDD

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.