REVIEW 3 major objections 4 minor 136 references
Hyperbolic Deep Learning for Foundation Models: A Survey
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This survey argues that Euclidean geometry is a fundamental bottleneck for foundation models and that hyperbolic space, with its exponential volume growth, is a mathematically grounded alternative that improves reasoning, generalization…
desk verdict A well-organized survey whose taxonomy is worth keeping, but the abstract overstates what the cited comparisons actually establish, since geometry is never isolated from co-introduced architecture changes. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is hyperbolic space $H^{K,n}$, a manifold of constant negative curvature $-K$, used in the Poincaré ball model and the Lorentz hyperboloid model. Its defining property, exponential volume growth with distance, is what lets trees and power-law distributions embed with low distortion, and this property is the mathematical engine of the survey's argument. The machinery that carries the argument into practice is the catalog of operations built on that space: tangent-space maps using $\exp$ and $\log$, Möbius addition and multiplication, fully hyperbolic Lorentz transformations, curvature-adaptive transformations, hyperbolic midpoints for attention, residual connection schemes, normalization layers, hyperbolic rotary positional encodings, and hyperbolic contrastive and entailment losses. These replace Euclidean layers one by one, and the survey uses them to classify each foundation model as hybrid, tangent-space, or fully hyperbolic.
What would settle it
A controlled head-to-head in which a hyperbolic LLM and a Euclidean LLM are trained from scratch on identical data with matched parameter count, compute budget, and hyperparameter tuning—and the hyperbolic model fails to beat the Euclidean model on hierarchy-rich reasoning benchmarks—would undercut the central claim. A simpler audit is to check whether the cited comparisons hold up when training budgets and search effort are equalized.
Extended reading notes
Core claim
The paper's central claim, assembled from the literature it reviews, is that the Euclidean inductive bias is a fundamental limitation of current foundation models and that hyperbolic geometry is a mathematically grounded alternative. Its core discovery is a working toolkit: operations that let Transformer, vision, and multimodal architectures run directly on hyperbolic manifolds, including hyperbolic attention, normalization, residual connections, positional encodings, and hyperbolic versions of contrastive and entailment losses. The survey reports that models built with this toolkit outperform Euclidean counterparts on hierarchy-rich tasks, and that the progression from hybrid designs to fully hyperbolic designs is correlated with stronger capture of structured information. It also reports the first hyperbolic LLMs at billion-parameter scale, fully hyperbolic vision transformers, and fully hyperbolic CLIP-style models.
Load-bearing premise
The survey's case depends on the reported benchmark results for the hyperbolic models it cites being obtained under fair comparisons with Euclidean baselines at matched scale, compute, and hyperparameter care.
Editorial extensions
If this is right
- Transformer-based models can swap Euclidean attention, normalization, residuals, and positional encoding for hyperbolic counterparts while keeping the overall architecture intact.
- Hyperbolic LLMs at billion-parameter scale can match or beat Euclidean LLMs on several benchmarks while using latent key-value caching to shrink generation-time memory.
- Vision-language models trained with hyperbolic contrastive and entailment losses can represent partial-order semantics, improving zero-shot generalization beyond Euclidean and spherical counterparts.
- Hyperbolic linear attention gives graph Transformers a route to large graphs, where quadratic-time hyperbolic attention previously failed.
- Because hyperbolic embeddings need fewer dimensions for the same representational quality, model scaling can become more parameter-efficient at high dimensions.
Reading between the lines
- A direct extension the survey leaves implicit: if hyperbolic embeddings truly compress hierarchies into few dimensions, retrieval-augmented generation and knowledge-graph search could be rebuilt around hyperbolic nearest-neighbor search, and the paper's proposed hyperbolic RAG is the natural testbed.
- The survey documents that fully hyperbolic pre-training is still at a fraction of Euclidean scale; a fair inference is that the strongest test of the central claim has not yet been run, and could go either way.
- The same low-distortion property that serves language hierarchies should transfer to structured biological data such as protein interaction networks and single-cell taxonomies, which the survey mentions but does not develop.
- The survey's own progression from hybrid to fully hyperbolic Transformers hints that the endgame is a fully hyperbolic stack—including Riemannian optimizers and attention kernels—rather than hyperbolic embeddings bolted onto Euclidean backbones.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey argues that hyperbolic geometry is a better inductive bias than Euclidean geometry for foundation models, because hyperbolic spaces can embed hierarchical and power-law data with low distortion in fewer dimensions. It introduces the mathematical background, catalogues hyperbolic building blocks (Eqs. 1–12), and organizes existing hyperbolic transformers/LLMs, vision models, and multimodal models (Sec. 4), closing with challenges (Sec. 5). The paper's central empirical assertion, stated in the abstract and repeated in Sec. 4.1, is that recent hyperbolic models outperform Euclidean counterparts in LLM reasoning, VLM generalization, and cross-modal alignment while maintaining parameter efficiency.
Significance. If the central claim were supported by controlled comparisons, the survey would be a valuable roadmap for a promising direction: the technical catalogue of tangent-space, fully-hyperbolic, and refined operations (Sec. 3) is useful, and the taxonomy by modality and geometry mode (Table 1) is clear. The paper also honestly lists open problems, including the lack of large-scale hyperbolic pretraining and Riemannian training tools. However, the comparative evidence for geometric superiority is not yet established: the cited successes change multiple architectural components at once, and the strongest results come from the authors' own papers without benchmark tables or independent replication. The survey's value is therefore primarily as an organized review, not as a demonstration that hyperbolic geometry is the cause of the reported gains.
major comments (3)
- [Abstract; §4.1] The claim that hyperbolic geometry is a superior inductive bias is load-bearing, but the cited comparisons do not isolate geometry as the active ingredient. HELM (ref [47]) couples hyperbolic modules with a new Multi-Head Latent Attention and Mixture-of-Curvature-Experts; Hypformer (ref [115]) introduces hyperbolic linear attention plus readjustment/refinement blocks; LResNet (ref [50]) changes the residual connection; H-BLIP-2 (§4.3) changes the alignment loss and adds a stabilization term. In none of these cases does the Euclidean baseline receive the non-geometric counterpart of the new component with only the manifold removed. The reported gains are therefore equally compatible with the alternative explanation that the co-introduced architecture, not negative curvature, drives the improvement. The survey should either present controlled ablations or explicitly qualify the central claim.
- [§4.1–§4.3] The statement in §4.1 that 'HELM models outperformed Euclidean LLMs in billions-of-parameters scale on several benchmarks' cites only the authors' preprint [47] and is presented without any benchmark table, metric values, error bars, or training/evaluation budgets. Similar author-reported claims appear for H-BERT, Hypformer, HypLoRA, HyperCore/LViT/L-CLIP, and H-BLIP-2 in §§4.1–4.3. Since these results are the primary evidence for the survey's central claim, a survey should either include a consolidated results table with model sizes, datasets, and protocols, or flag these as unverified author-reported results and note the absence of independent replication.
- [§5] The paper's own assessment concedes that hyperbolic LLM pretraining is still at about one billion parameters, far below Euclidean-scale pretraining, that no hyperbolic vision foundation model has been pretrained, and that standard Riemannian tools such as a full AdamW equivalent and FlashAttention support are missing. These concessions directly undercut the abstract's 'while maintaining parameter efficiency' and the implied claim that hyperbolic geometry already enhances foundation models at scale. The abstract and conclusion should be reworded to present hyperbolic geometry as a promising but not yet established alternative, consistent with Section 5.
minor comments (4)
- [§3.1, Eq. (11)] Equation (11) contains a factor sqrt(K2/K2), which is identically 1; this is likely a typo for a curvature ratio involving K1 and K2, and should be corrected.
- [Throughout] There are several typos and inconsistent notations, including 'Poincaré aall model' (Appendix A), 'tengant' (Sec. 2.1), 'curvaure' (Sec. 3.2), 'Attnetion' (Sec. 3.1), and 'MEUR' (Sec. 4.3, should be MERU).
- [§5.4] The claim that 'The Hyp-GraphRAG model from HyperCore [49] provides a proof of concept' is not supported by any details or experimental result; either add a concrete reference to a result or remove the claim.
- [Figure 1] The caption says 'The red line shows the same geodesic under an isometry' but the two panels display different models of hyperbolic space; the caption should clarify what exactly is being compared.
Circularity Check
Survey's LLM/Transformer performance claims lean on self-citations (HELM, Hypformer), but independent works give the central thesis non-circular support.
-
self citation load bearing
[Section 4.1, Hyperbolic LLMs (paragraph on HELM, ref [47])]
"When trained on the same datasets, HELM models outperformed Euclidean LLMs in billions-of-parameters scale on several benchmarks."
The survey's central claim that hyperbolic geometry improves LLM capabilities at scale is supported in this passage solely by citation [47], HELM, whose author list includes the present survey's authors (He, Madhu, Yang, Ying). The survey does not reproduce or independently verify the benchmark comparison; the load-bearing evidence for the abstract's promise of 'improving LLMs' complex reasoning ability ... while maintaining parameter efficiency' is thus the authors' own prior work, not an external, machine-checked, or independently reproduced result.
-
self citation load bearing
[Section 4.1, Hyperbolic Transformers (paragraph on Hypformer, ref [115])]
"Whether incorporating the GNN or not, Hypformer models outperformed Euclidean Transformer models on a variety of graph, image, and text tasks."
Ref [115] (Hypformer) shares two authors with the present survey (Menglin Yang and Rex Ying). This sentence is the survey's primary evidence that fully hyperbolic Transformers beat Euclidean baselines across modalities, which is load-bearing for the abstract's claim that hyperbolic advances 'enhance foundation models'. No independent benchmark, external reproduction, or falsifiable analysis is cited in the survey for this specific claim. The circularity is partial because independent works (HAN, FNN, H-BERT) also appear in the same section.
full rationale
This paper is a literature survey rather than a derivation chain, so there is no equation-level circularity of the form where an output is identical to an input by construction. The mathematical background on low-distortion tree and power-law embeddings (refs [98,100]) is independent and non-circular. The central empirical thesis that hyperbolic geometry benefits foundation models is supported by many cited primary works, several of which (HELM [47], Hypformer [115], HyperCore [49], LResNet [50], HypLoRA [113]) overlap with the present authors. The two strongest scale-level claims for hyperbolic LLMs and fully hyperbolic Transformers rest on self-citations with no independent verification inside the survey, which is a load-bearing self-citation pattern rather than normal background citation. However, the survey also covers independent works (HAN, FNN, H-BERT, MERU, Hyp-ViT, H-BLIP-2, HyCoCLIP) that support the same general thesis, and Section 5 candidly concedes that hyperbolic pretraining remains far below Euclidean scale in model size and dataset size. The central claim therefore retains independent content and is not forced by the self-citation chain. The reviewer concern about confounded comparisons (multiple architectural components changing at once) is a legitimate experimental-design and correctness risk, but it is not circularity under the definitions used here.
Assumptions & free parameters
assumptions (4)
- standard math Hyperbolic space allows low-distortion embedding of trees and power-law distributions (e.g., Sarkar 2011; Sala et al. 2018).
- domain assumption Real-world data in language, vision, and graphs is often hierarchical and scale-free.
- domain assumption The reported performance gains of hyperbolic models over Euclidean baselines are reliable and not due to confounding factors.
- domain assumption Euclidean foundation models have fundamental limitations that are not addressable by better Euclidean training.
Cite this review
Pith. "Pith review of Hyperbolic Deep Learning for Foundation Models: A Survey." pith.science (2026). https://pith.science/paper/LF3VHLHE
@misc{pith2026250717787,
author = {Pith},
title = {Pith review of: Hyperbolic Deep Learning for Foundation Models: A Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/LF3VHLHE}},
note = {Machine review of arXiv:2507.17787}
}
read the original abstract
Foundation models pre-trained on massive datasets, including large language models (LLMs), vision-language models (VLMs), and large multimodal models, have demonstrated remarkable success in diverse downstream tasks. However, recent studies have shown fundamental limitations of these models: (1) limited representational capacity, (2) lower adaptability, and (3) diminishing scalability. These shortcomings raise a critical question: is Euclidean geometry truly the optimal inductive bias for all foundation models, or could incorporating alternative geometric spaces enable models to better align with the intrinsic structure of real-world data and improve reasoning processes? Hyperbolic spaces, a class of non-Euclidean manifolds characterized by exponential volume growth with respect to distance, offer a mathematically grounded solution. These spaces enable low-distortion embeddings of hierarchical structures (e.g., trees, taxonomies) and power-law distributions with substantially fewer dimensions compared to Euclidean counterparts. Recent advances have leveraged these properties to enhance foundation models, including improving LLMs' complex reasoning ability, VLMs' zero-shot generalization, and cross-modal semantic alignment, while maintaining parameter efficiency. This paper provides a comprehensive review of hyperbolic neural networks and their recent development for foundation models. We further outline key challenges and research directions to advance the field.
Figures
Reference graph
Works this paper leans on
-
[47]
Neil He, Rishabh Anand, Hiren Madhu, Ali Maatouk, Smita Krishnaswamy, Lean- dros Tassiulas, Menglin Yang, and Rex Ying. 2025. HELM: Hyperbolic Large Lan- guage Models via Mixture-of-Curvature Experts. arXiv preprint arXiv:2505.24722 (2025)
arXiv 2025
-
[115]
Menglin Yang, Harshit Verma, Delvin Ce Zhang, Jiahong Liu, Irwin King, and Rex Ying. 2024. Hypformer: Exploring efficient transformer fully in hyperbolic space. In KDD. 3770–3781
work page 2024
-
[50]
Neil He, Menglin Yang, and Rex Ying. 2025. Lorentzian Residual Neural Net- works. In KDD
2025
-
[1]
Gregorio Alanis-Lobato, Pablo Mier, and Miguel Andrade-Navarro. 2018. The latent geometry of the human protein interaction network. Bioinformatics 34, 16 (2018), 2826–2834
2018
-
[2]
Muhammad Awais, Muzammal Naseer, Salman Khan, Rao Muhammad Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang, and Fahad Shahbaz Khan
-
[3]
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer Normal- ization. arXiv arXiv:1607.06450 (2016)
arXiv 2016
-
[4]
Yushi Bai, Zhitao Ying, Hongyu Ren, and Jure Leskovec. 2021. Modeling hetero- geneous hierarchies with relation-specific hyperbolic cones. NeurIPS 34 (2021), 12316–12327
2021
-
[5]
Federico Barbero, Alex Vitvitskyi, Christos Perivolaropoulos, Razvan Pascanu, and Petar Veličković. 2025. Round and Round We Go! What makes Rotary Positional Encodings useful? arXiv:2410.06205 (2025)
arXiv 2025
Show all 136 references
-
[6]
Ahmad Bdeir, Kristian Schwethelm, and Niels Landwehr. 2024. Fully Hyperbolic Convolutional Neural Networks for Computer Vision. In ICLR
2024
-
[7]
Nithya Bhasker, Hattie Chung, Louis Boucherie, Vladislav Kim, Stefanie Speidel, and Melanie Weber. 2024. Contrastive poincaré maps for single-cell data analysis. In ICLR 2024 Workshop on Machine Learning for Genomics Explorations
2024
-
[8]
Rishi Bommasani and et al. 2021. On the Opportunities and Risks of Foundation Models. arXiv preprint arXiv:2108.07258 (2021)
2021 arXiv
-
[9]
Joey Bose, Ariella Smofsky, Renjie Liao, Prakash Panangaden, and Will Hamilton
-
[10]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. NeurIPS 33 (2020), 1877–1901
2020
-
[11]
Davide Caffagni, Federico Cocchi, Luca Barsellotti, Nicholas Moratelli, Sara Sarto, Lorenzo Baraldi, Marcella Cornia, and Rita Cucchiara. 2024. The revolu- tion of multimodal large language models: a survey. arXiv:2402.12451 (2024)
2024
-
[12]
Hanqun Cao, Cheng Tan, Zhangyang Gao, Yilun Xu, Guangyong Chen, Pheng- Ann Heng, and Stan Z Li. 2024. A survey on generative diffusion models. TKDE (2024)
2024
-
[13]
Ines Chami, Zhitao Ying, Christopher Ré, and Jure Leskovec. 2019. Hyperbolic graph convolutional neural networks. In NeurIPS. 4868–4879
2019
-
[14]
Jiayang Chen, Zhihang Hu, Siqi Sun, Qingxiong Tan, Yixuan Wang, Qinze Yu, Licheng Zong, Liang Hong, Jin Xiao, Tao Shen, et al. 2022. Interpretable RNA foundation model from unannotated data for highly accurate RNA structure and function predictions. arXiv:2204.00300 (2022)
2022 arXiv
-
[15]
Ming Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding, and Yaliang Li. 2020. Simple and deep graph convolutional networks. In ICML. PMLR, 1725–1735
2020
-
[16]
Ricky T. Q. Chen and Yaron Lipman. 2024. Flow Matching on General Geometries. In ICLR
2024
-
[17]
Weize Chen, Xu Han, Yankai Lin, Kaichen He, Ruobing Xie, Jie Zhou, and Zhiyuan Liu. 2024. Hyperbolic Pre-Trained Language Model. IEEE TASLP 32 (2024)
2024
-
[18]
Weize Chen, Xu Han, Yankai Lin, Hexu Zhao, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2021. Fully Hyperbolic Neural Networks. arXiv:2105.14686 (2021)
2021 arXiv
-
[19]
Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. 2020. Improved base- lines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, (2020)
2020 arXiv
-
[20]
Yankai Chen, Menglin Yang, Yingxue Zhang, Mengchen Zhao, Ziqiao Meng, Jianye Hao, and Irwin King. 2022. Modeling Scale-free Graphs for Knowledge- aware Recommendation. WSDM (2022)
2022
-
[21]
Jeffrey Cheng, Marc Marone, Orion Weller, Dawn Lawrie, Daniel Khashabi, and Benjamin Van Durme. 2024. Dated data: Tracing knowledge cutoffs in large language models. arXiv:2403.12958 (2024)
2024 arXiv
-
[22]
Andy Coenen, Emily Reif, Ann Yuan, Been Kim, Adam Pearce, Fernanda Viégas, and Martin Wattenberg. 2019. Visualizing and Measuring the Geometry of BERT. NuerIPS (2019)
2019
-
[23]
Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah
-
[24]
Haotian Cui, Chloe Wang, Hassaan Maan, Kuan Pang, Fengning Luo, Nan Duan, and Bo Wang. 2024. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nature Methods 21, 8 (2024), 1470–1480
2024
-
[25]
Jindou Dai, Yuwei Wu, Zhi Gao, and Yunde Jia. 2021. A Hyperbolic-to- Hyperbolic Graph Convolutional Network. arXiv:2104.06942 (2021), 154–163
2021 arXiv
-
[26]
Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. In NeurIPS
2022
-
[27]
Valentin De Bortoli, Emile Mathieu, Michael Hutchinson, James Thornton, Yee Whye Teh, and Arnaud Doucet. 2022. Riemannian Score-Based Generative Modelling. In NeurIPS
2022
-
[28]
DeepSeek-AI. 2024. DeepSeek-V3 Technical Report. arXiv:2412.19437 (2024)
2024 arXiv
-
[29]
Karan Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson, and Shan- mukha Ramakrishna Vedantam. 2023. Hyperbolic image-text representations. In ICML. PMLR, 7694–7731
2023
-
[30]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805 (2018)
2018 arXiv
-
[31]
Jiarui Ding and Aviv Regev. 2021. Deep generative model embedding of single- cell RNA-Seq profiles on hyperspheres and hyperbolic spaces. Nature communi- cations 12, 1 (2021), 2554
2021
-
[32]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al . 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11...
2020 arXiv
-
[33]
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, and Steven Truitt. 2024. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv:2404.16130 (2024)
2024 arXiv
-
[34]
Aleksandr Ermolov, Leyla Mirvakhabova, Valentin Khrulkov, Nicu Sebe, and Ivan Oseledets. 2022. Hyperbolic vision transformers: Combining improvements in metric learning. In CVPR. 7409–7419
2022
-
[35]
Jacob Fein-Ashley, Ethan Feng, and Minh Pham. 2024. HVT: A Comprehensive Vision Framework for Learning in Non-Euclidean Space. arXiv:2409.16897 (2024)
2024 arXiv
-
[36]
Xingcheng Fu, Yisen Gao, Yuecen Wei, Qingyun Sun, Hao Peng, Jianxin Li, and Xianxian Li. 2024. Hyperbolic Geometric Latent Diffusion Model for Graph Generation. ICML (2024)
2024
-
[37]
Octavian Ganea, Gary Bécigneul, and Thomas Hofmann. 2018. Hyperbolic entailment cones for learning hierarchical embeddings. In ICML. PMLR, 1646– 1655
2018
-
[38]
Octavian Ganea, Gary Bécigneul, and Thomas Hofmann. 2018. Hyperbolic neural networks. In NeurIPS. 5345–5355
2018
-
[39]
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv:2312.10997 2 (2023)
2023 arXiv
-
[40]
Songwei Ge, Shlok Mishra, Simon Kornblith, Chun-Liang Li, and David Jacobs
-
[41]
Songwei Ge, Shlok Kumar Mishra, Simon Kornblith, Chun-Liang Li, and David Jacobs. 2022. Hyperbolic Contrastive Learning for Visual Representations be- yond Objects. ArXiv abs/2212.00653 (2022)
2022 arXiv
-
[42]
John Clifford Gower. 1985. Properties of Euclidean and non-Euclidean distance matrices. Linear algebra and its applications 67 (1985), 81–97
1985
-
[43]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The Llama 3 Herd of Models.arXiv:2407.21783 (2024)
2024 arXiv
-
[44]
Hyperbolic contrastive learning for visual representations beyond objects. In CVPR. 6840–6849
-
[45]
Zihao Guo, Qingyun Sun, Haonan Yuan, Xingcheng Fu, Min Zhou, Yisen Gao, and Jianxin Li. 2025. GraphMoRE: Mitigating Topological Heterogeneity via Mixture of Riemannian Experts. In AAAI
2025
-
[46]
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Mo- mentum contrast for unsupervised visual repre- sentation learning. In CVPR
2020
-
[48]
Caglar Gulcehre, Misha Denil, Mateusz Malinowski, Ali Razavi, Razvan Pas- canu, Karl Moritz Hermann, Peter Battaglia, Victor Bapst, David Raposo, Adam Santoro, et al. 2019. Hyperbolic attention networks. In ICLR
2019
-
[49]
Neil He, Menglin Yang, and Rex Ying. 2025. HyperCore: The Core Framework for Building Hyperbolic Foundation Models with Comprehensive Modules. arXiv:2504.08912 (2025)
2025 arXiv
-
[51]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising Diffusion Proba- bilistic Models. In NeurIPS
2020
-
[52]
Neil He, Jiahong Liu, Buze Zhang, Ngoc Bui, Ali Maatouk, Menglin Yang, Irwin King, Melanie Weber, and Rex Ying. 2025. Position: Beyond Euclidean–Foundation Models Should Embrace Non-Euclidean Geometries. arXiv:2504.08896 (2025)
2025
-
[53]
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In ICLR
2022
-
[54]
Chin-Wei Huang, Milad Aghajohari, Avishek Joey Bose, Prakash Panangaden, and Aaron Courville. 2022. Riemannian Diffusion Models. In NeurIPS
2022
-
[55]
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2025. KDD ’25, August 3–7, 2025, Toronto, ON, Canada Neil He, Hiren Madhu, Ngoc Bui, Menglin Yang, and Rex Ying A survey on hallucinatio...
2025
-
[56]
Jordan Hoffmann and et al. 2022. Training compute-optimal large language models. In NeurIPS. Red Hook, NY, USA, Article 2176, 15 pages
2022
-
[57]
Jaehyeong Jo and Hwang Sung Ju Lee, Seul and. 2022. Score-based Generative Modeling of Graphs via the System of Stochastic Differential Equations. In ICML
2022
-
[58]
Ho, Percy Liang, and Arvind Narayanan
Sayash Kapoor, Rishi Bommasani, Kevin Klyman, Shayne Longpre, Ashwin Ramaswami, Peter Cihon, Aspen Hopkins, Kevin Bankston, Stella Biderman, Miranda Bogen, Rumman Chowdhury, Alex Engler, Peter Henderson, Yacine Jer- nite, Seth Lazar, Stefano Maffulli, Alondra Nelson, Joelle Pi...
2023 arXiv
-
[59]
W Sean Kennedy, Onuttom Narayan, and Iraj Saniee. 2013. On the hyperbolicity of large-scale networks. arXiv:1307.0031 (2013)
2013 arXiv
-
[60]
Sergey Ioffe and Christian Szegedy. 2015. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. InICML. 448–456
2015
-
[61]
Minseon Kim, Jihoon Tack, and Sung Ju Hwang. 2020. Adversarial Self- Supervised Contrastive Learning. In NeurIPS
2020
-
[62]
Kingma and Max Welling
Diederik P. Kingma and Max Welling. 2019. An Introduction to Variational Autoencoders. Foundations and Trends in Machine Learning, Vol. 12. 307–392 pages
2019
-
[63]
Max Kochurov, Rasul Karimov, and Sergei Kozlukov. 2020. Geoopt: Riemannian Optimization in PyTorch. arXiv:2005.02819 (2020)
2020 arXiv
-
[64]
Bobak Kiani, Jason Wang, and Melanie Weber. 2024. Hardness of Learning Neural Networks under the Manifold Hypothesis. In NeurIPS
2024
-
[65]
Marc Law, Renjie Liao, Jake Snell, and Richard Zemel. 2019. Lorentzian distance learning for hyperbolic representations. In ICML. PMLR, 3672–3681
2019
-
[66]
Diego Lazcano, Nicolás Fredes, and Werner Creixell. 2021. Hyperbolic Genera- tive Adversarial Network. IEEE Access 9 (2021), 96309–96320
2021
-
[67]
Matthew Le, Stephen Roller, Laetitia Papaxanthos, Douwe Kiela, and Maximilian Nickel. 2019. Inferring Concept Hierarchies from Text Corpora via Hyperbolic Embeddings. In ACL. 3231–3241
2019
-
[68]
Dmitri Krioukov, Fragkiskos Papadopoulos, Maksim Kitsak, Amin Vahdat, and Marián Boguná. 2010. Hyperbolic geometry of complex networks. Physical Review E—Statistical, Nonlinear, and Soft Matter Physics 82, 3 (2010), 036106
2010
-
[69]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rock- täschel, et al. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In NeurIPS. 9459–9774
2020
-
[70]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Boot- strapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In ICML
2023
-
[71]
Lingxiao Li, Yi Zhang, and Shuhui Wang. 2023. The Euclidean Space is Evil: Hyperbolic Attribute Editing for Few-shot Image Generation. In ICCV
2023
-
[72]
Holden Lee, Jianfeng Lu, and Yixin Tan. 2022. Convergence for score-based generative modeling with polynomial complexity. In NeurIPS
2022
-
[73]
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. 2022. Flow Matching for Generative Modeling. arXiv:2210.02747 (2022)
2022 arXiv
-
[74]
Jiahong Liu, Menglin Yang, Min Zhou, Shanshan Feng, and Philippe Fournier- Viger. 2022. Enhancing Hyperbolic Graph Embeddings via Contrastive Learning. In NeurIPS 2nd SSL workshop
2022
-
[75]
Qi Liu, Maximilian Nickel, and Douwe Kiela. 2019. Hyperbolic graph neural networks. In NeurIPS. 8230–8241
2019
-
[76]
Nathan Linial, Eran London, and Yuri Rabinovich. 1995. The geometry of graphs and some of its algorithmic applications. Combinatorica 15, 2 (1995), 215–245
1995
-
[77]
Aaron Lou, Isay Katsman, Qingxuan Jiang, Serge Belongie, Ser-Nam Lim, and Christopher De Sa. 2020. Differentiating through the Fréchet Mean. In ICML. 6393–6403
2020
-
[78]
Qiyao Ma, Menglin Yang, Mingxuan Ju, Tong Zhao, Neil Shah, and Rex Ying
-
[79]
Paolo Mandica, Luca Franco, Konstantinos Kallidromitis, Suzanne Petryk, and Fabio Galasso. 2024. Hyperbolic Learning with Multimodal Large Language Models. arXiv:2408.05097 (2024)
2024 arXiv
-
[80]
Aaron Lou, Isay Katsman, Qingxuan Jiang, Serge Belongie, Ser-Nam Lim, and Christopher De Sa. 2020. Differentiating through the Fr \’echet Mean. arXiv:2003.00335 (2020)
2020 arXiv
-
[81]
Pascal Mettes, Mina Ghadimi Atigh, Martin Keller-Ressel, Jeffrey Gu, and Serena Yeung. 2024. Hyperbolic deep learning in computer vision: A survey. Interna- tional Journal of Computer Vision (2024), 1–25
2024
-
[82]
Nina Miolane, Nicolas Guigui, Alice Le Brigant, Johan Mathe, Benjamin Hou, Yann Thanwerdas, Stefan Heyder, Olivier Peltre, Niklas Koep, Hadi Zaatiti, Hatem Hajri, Yann Cabanes, Thomas Gerald, Paul Chauchat, Christian Shew- make, Daniel Brooks, Bernhard Kainz, Claire Donnat, Su...
2020
-
[83]
Devon Myers, Rami Mohawesh, Venkata Ishwarya Chellaboina, Anantha Lak- shmi Sathvik, Praveen Venkatesh, Yi-Hui Ho, Hanna Henshaw, Muna Al- hawawreh, David Berdik, and Yaser Jararweh. 2024. Foundation and large language models: fundamentals, challenges, opportunities, and socia...
2024
-
[84]
Yoshihiro Nagano, Shoichiro Yamaguchi, Yasuhiro Fujita, and Masanori Koyama
-
[85]
Maddison, Ryota Tomioka, and Yee Whye Teh
Emile Mathieu, Charline Le Lan, Chris J. Maddison, Ryota Tomioka, and Yee Whye Teh. 2019. Continuous Hierarchical Representations with Poincaré Variational Auto-Encoders. In NeurIPS
2019
-
[86]
Maximillian Nickel and Douwe Kiela. 2018. Learning Continuous Hierarchies in the Lorentz Model of Hyperbolic Geometry. In ICML. 3779–3788
2018
-
[87]
Keiron O’Shea and Ryan Nash. 2015. An Introduction to Convolutional Neural Networks. arXiv:1511.08458 (2015)
2015 arXiv
-
[88]
Avik Pal, Max van Spengler, Guido Maria D’Amely di Melendugno, Alessandro Flaborea, Fabio Galasso, and Pascal Mettes. 2025. Compositional Entailment Learning for Hyperbolic Vision-Language Models. ICLR (2025)
2025
-
[89]
Fragkiskos Papadopoulos, Dmitri Krioukov, Marián Boguná, and Amin Vah- dat. 2010. Greedy forwarding in dynamic scale-free networks embedded in hyperbolic metric spaces. In 2010 Proceedings IEEE Infocom . IEEE, 1–9
2010
-
[90]
Wei Peng, Tuomas Varanka, Abdelrahman Mostafa, Henglin Shi, and Guoying Zhao. 2021. Hyperbolic deep neural networks: A survey. TPAMI (2021)
2021
-
[91]
Maximillian Nickel and Douwe Kiela. 2017. Poincaré embeddings for learning hierarchical representations. In NeurIPS. 6338–6347
2017
-
[92]
Aleksandar Poleksic. 2023. Hyperbolic matrix factorization improves prediction of drug-target associations. Scientific Reports 13, 1 (2023), 959
2023
-
[94]
Eric Qu and Dongmian Zou. 2022. Lorentzian fully hyperbolic generative adversarial network. arXiv:2201.12825 (2022)
2022 arXiv
-
[95]
Radford, Kim Jong Wook Alec, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. arXiv:2103.000...
2021 arXiv
-
[96]
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deep- Speed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters. In KDD. 3505–3506
2020
-
[97]
Xavier Pennec. 2006. Intrinsic Statistics on Riemannian Manifolds: Basic Tools for Geometric Measurements. JMIV 25, 1 (2006), 127–154
2006
-
[98]
Frederic Sala, Chris De Sa, Albert Gu, and Christopher Re. 2018. Representation Tradeoffs for Hyperbolic Embeddings. In ICML. 4460–4469
2018
-
[99]
Lillemark, Christian Shewmake, Abby Bertics, Xavier Pennec, and Nina Miolane
Sophia Sanborn, Johan Mathe, Mathilde Papillon, Domas Buracas, Hansen J. Lillemark, Christian Shewmake, Abby Bertics, Xavier Pennec, and Nina Miolane
-
[100]
Rik Sarkar. 2011. Low distortion delaunay embedding of trees in hyperbolic plane. In International Symposium on Graph Drawing . Springer, 355–366
2011
-
[101]
Ryohei Shimizu, Yusuke Mukuta, and Tatsuya Harada. 2020. Hyperbolic Neural Networks++. In ICLR
2020
-
[102]
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Ste- fano Ermon, and Ben Poole. 2021. Score-Based Generative Modeling through Stochastic Differential Equations. In ICLR
2021
-
[103]
Salem Said, Lionel Bombrun, and Yannick Berthoumieu. 2014. New Riemannian Priors on the Univariate Normal Model. Entropy 16, 7 (2014), 4015–4031
2014
-
[104]
Xunzhu Tang, Saad Ezzini, Haoye Tian, Yewei Song, Jacques Klein, Tegawende F Bissyande, et al. 2023. Hyperbolic code retrieval: a novel approach for efficient code search using hyperbolic space embeddings. arXiv:2308.15234 (2023)
2023 arXiv
-
[105]
Alexandru Tifrea, Gary Bécigneul, and Octavian-Eugen Ganea. 2018. Poincaré glove: Hyperbolic word embeddings. arXiv:1810.06546 (2018)
2018 arXiv
-
[106]
arXiv:2407.09468 (2024)
Beyond Euclid: An Illustrated Guide to Modern Machine Learning with Geometric, Topological, and Algebraic Structures. arXiv:2407.09468 (2024)
2024 arXiv
-
[107]
Max van Spengler, Erwin Berkhout, and Pascal Mettes. 2023. Poincaré ResNet. CVPR (2023)
2023
-
[108]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NeurIPS. 5998–6008
2017
-
[109]
Amelia Villegas-Morcillo, Stavros Makrodimitris, Roeland C H J van Ham, Angel M Gomez, Victoria Sanchez, and Marcel J T Reinders. 2020. Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function. Bioinformatics 37, ...
2020
-
[110]
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2021. RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv:2104.09864 (2021)
2021 arXiv
-
[111]
Lingfeng Wen, Xuan Tang, Mingjie Ouyang, Xiangxiang Shen, Jian Yang, Daxin Zhu, Mingsong Chen, and Xian Wei. 2024. Hyperbolic Graph Diffusion Model. In AAAI. Hyperbolic Deep Learning for Foundation Models: A Survey KDD ’25, August 3–7, 2025, Toronto, ON, Canada
2024
-
[112]
Jiahao Xie, Xiaohang Zhan, Ziwei Liu, Yew Soon Ong, and Chen Change Loy
-
[113]
Abraham Albert Ungar. 2008. A gyrovector space approach to hyperbolic geometry. Synthesis Lectures on Mathematics and Statistics 1, 1 (2008), 1–194
2008
-
[114]
Menglin Yang, Zhihao Li, Min Zhou, Jiahong Liu, and Irwin King. 2022. Hicf: Hyperbolic informative collaborative filtering. In KDD. 2212–2221
2022
-
[116]
Menglin Yang, Min Zhou, Marcus Kalander, Zengfeng Huang, and Irwin King
-
[117]
Xinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong, and Lei Li. 2021. Dense contrastive learning for self-supervised visual pre-training. In CVPR
2021
-
[118]
Menglin Yang, Min Zhou, Jiahong Liu, Defu Lian, and Irwin King. 2022. HRCF: Enhancing Collaborative Filtering via Hyperbolic Geometric Regularization. In WebConf
2022
-
[119]
Menglin Yang, Min Zhou, Hui Xiong, and Irwin King. 2022. Hyperbolic Temporal Network Embedding. TKDE (2022)
2022
-
[120]
Yamada, and Ziming Zhang
Yun Yue, Fangzhou Lin, Kazunori D. Yamada, and Ziming Zhang. 2023. Hyper- bolic Contrastive Learning. arXiv:2302.01409 (2023)
2023 arXiv
-
[121]
Menglin Yang, Aosong Feng, Bo Xiong, Jihong Liu, Irwin King, and Rex Ying
-
[122]
ICML LLM Cognition Workshop (2024)
Hyperbolic Fine-tuning for Large Language Models. ICML LLM Cognition Workshop (2024)
2024
-
[123]
Yiding Zhang, Xiao Wang, Chuan Shi, Nian Liu, and Guojie Song. 2021. Lorentzian Graph Convolutional Networks. In WebConf. 1249–1261
2021
-
[124]
Yifei Zhang, Hao Zhu, Menglin Yang, Jiahong Liu, Rex Ying, Irwin King, and Piotr Koniusz. 2025. Understanding and mitigating hyperbolic dimensional collapse in graph contrastive learning. In KDD. 1984–1995
2025
-
[125]
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv:2303.18223 1, 2 (2023)
2023 arXiv
-
[126]
Discrete-time Temporal Network Embedding via Implicit Hierarchical Learning in Hyperbolic Space. In KDD. 1975–1985
1975
-
[127]
Menglin Yang, Min Zhou, Zhihao Li, Jiahong Liu, Lujia Pan, Hui Xiong, and Irwin King. 2022. Hyperbolic graph neural networks: a review of methods and applications. arXiv:2202.13852 (2022)
2022 arXiv
-
[131]
Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu. 2024. Vision-Language Models for Vision Tasks: A Survey. IEEE TPAMI 46, 8 (2024), 5625–5644
2024
-
[132]
Yiding Zhang, Xiao Wang, Chuan Shi, Xunqiang Jiang, and Yanfang Fanny Ye
-
[133]
TBD (2021)
Hyperbolic graph attention network. TBD (2021)
2021
-
[137]
Shichao Zhu, Shirui Pan, Chuan Zhou, Jia Wu, Yanan Cao, and Bin Wang. 2020. Graph Geometry Interaction Learning. In NeurIPS, Vol. 33. 7548–7558. A Hyperbolic Geometry and Additional Works A.1 Hyperbolic Geometry Lorentz model. An𝑛-dimensional Lorentz model is a Riemann- ian ma...
2020
-
[2019]
A wrapped normal distribution on hyperbolic space for gradient-based learning. In ICML. PMLR, 4693–4702
-
[2020]
Latent variable modelling with hyperbolic normalizing flows. In ICML. PMLR, 1045–1055
-
[2021]
In NeurIPS
Unsupervised object-level representation learning from scene images. In NeurIPS
-
[2023]
TPAMI (2023)
Diffusion models in vision: A survey. TPAMI (2023)
2023
-
[2024]
arXiv:2411.13865 (2024)
HARec: Hyperbolic graph-llm alignment for exploration and exploitation in recommender systems. arXiv:2411.13865 (2024)
2024 arXiv
-
[2025]
TPAMI (2025)
Foundation Models Defining a New Era in Vision: a Survey and Outlook. TPAMI (2025)
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.