Pith. sign in

REVIEW 3 major objections 4 minor 136 references

Hyperbolic Deep Learning for Foundation Models: A Survey

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This survey argues that Euclidean geometry is a fundamental bottleneck for foundation models and that hyperbolic space, with its exponential volume growth, is a mathematically grounded alternative that improves reasoning, generalization…

desk verdict A well-organized survey whose taxonomy is worth keeping, but the abstract overstates what the cited comparisons actually establish, since geometry is never isolated from co-introduced architecture changes. read the letter →

arxiv 2507.17787 v1 pith:LF3VHLHE submitted 2025-07-23 cs.LG cs.AI

classification cs.LGcs.AI
keywords HyperbolicgeometryFoundationmodelsLargelanguageVision-languageRepresentationlearningTransformerNon-EuclideanembeddingsHierarchicaldata
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Foundation models are almost always built in Euclidean space, where the geometry is flat. This survey argues that flatness is a bottleneck: real-world data—language syntax, taxonomies, social networks, protein and gene hierarchies—is tree-like and scale-free, and embedding such structure in Euclidean space takes many dimensions and still distorts it. The survey's central claim is that hyperbolic space, whose volume grows exponentially with distance, is a better inductive bias: it embeds hierarchies and power-law distributions in substantially fewer dimensions, and recent hyperbolic language, vision, and multimodal models built on this bias improve reasoning, zero-shot generalization, and cross-modal alignment while staying parameter-efficient. The contribution is a systematic map of the hyperbolic building blocks—attention, normalization, residuals, positional encodings, and contrastive and entailment losses—and of which current models are hybrid, tangent-space, or fully hyperbolic. If the claim is right, the geometry of large-scale AI becomes a choice worth making explicitly.

What carries the argument

The central object is hyperbolic space $H^{K,n}$, a manifold of constant negative curvature $-K$, used in the Poincaré ball model and the Lorentz hyperboloid model. Its defining property, exponential volume growth with distance, is what lets trees and power-law distributions embed with low distortion, and this property is the mathematical engine of the survey's argument. The machinery that carries the argument into practice is the catalog of operations built on that space: tangent-space maps using $\exp$ and $\log$, Möbius addition and multiplication, fully hyperbolic Lorentz transformations, curvature-adaptive transformations, hyperbolic midpoints for attention, residual connection schemes, normalization layers, hyperbolic rotary positional encodings, and hyperbolic contrastive and entailment losses. These replace Euclidean layers one by one, and the survey uses them to classify each foundation model as hybrid, tangent-space, or fully hyperbolic.

What would settle it

A controlled head-to-head in which a hyperbolic LLM and a Euclidean LLM are trained from scratch on identical data with matched parameter count, compute budget, and hyperparameter tuning—and the hyperbolic model fails to beat the Euclidean model on hierarchy-rich reasoning benchmarks—would undercut the central claim. A simpler audit is to check whether the cited comparisons hold up when training budgets and search effort are equalized.

Watch

Extended reading notes

Core claim

The paper's central claim, assembled from the literature it reviews, is that the Euclidean inductive bias is a fundamental limitation of current foundation models and that hyperbolic geometry is a mathematically grounded alternative. Its core discovery is a working toolkit: operations that let Transformer, vision, and multimodal architectures run directly on hyperbolic manifolds, including hyperbolic attention, normalization, residual connections, positional encodings, and hyperbolic versions of contrastive and entailment losses. The survey reports that models built with this toolkit outperform Euclidean counterparts on hierarchy-rich tasks, and that the progression from hybrid designs to fully hyperbolic designs is correlated with stronger capture of structured information. It also reports the first hyperbolic LLMs at billion-parameter scale, fully hyperbolic vision transformers, and fully hyperbolic CLIP-style models.

Load-bearing premise

The survey's case depends on the reported benchmark results for the hyperbolic models it cites being obtained under fair comparisons with Euclidean baselines at matched scale, compute, and hyperparameter care.

Editorial extensions

If this is right

  • Transformer-based models can swap Euclidean attention, normalization, residuals, and positional encoding for hyperbolic counterparts while keeping the overall architecture intact.
  • Hyperbolic LLMs at billion-parameter scale can match or beat Euclidean LLMs on several benchmarks while using latent key-value caching to shrink generation-time memory.
  • Vision-language models trained with hyperbolic contrastive and entailment losses can represent partial-order semantics, improving zero-shot generalization beyond Euclidean and spherical counterparts.
  • Hyperbolic linear attention gives graph Transformers a route to large graphs, where quadratic-time hyperbolic attention previously failed.
  • Because hyperbolic embeddings need fewer dimensions for the same representational quality, model scaling can become more parameter-efficient at high dimensions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct extension the survey leaves implicit: if hyperbolic embeddings truly compress hierarchies into few dimensions, retrieval-augmented generation and knowledge-graph search could be rebuilt around hyperbolic nearest-neighbor search, and the paper's proposed hyperbolic RAG is the natural testbed.
  • The survey documents that fully hyperbolic pre-training is still at a fraction of Euclidean scale; a fair inference is that the strongest test of the central claim has not yet been run, and could go either way.
  • The same low-distortion property that serves language hierarchies should transfer to structured biological data such as protein interaction networks and single-cell taxonomies, which the survey mentions but does not develop.
  • The survey's own progression from hybrid to fully hyperbolic Transformers hints that the endgame is a fully hyperbolic stack—including Riemannian optimizers and attention kernels—rather than hyperbolic embeddings bolted onto Euclidean backbones.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This survey argues that hyperbolic geometry is a better inductive bias than Euclidean geometry for foundation models, because hyperbolic spaces can embed hierarchical and power-law data with low distortion in fewer dimensions. It introduces the mathematical background, catalogues hyperbolic building blocks (Eqs. 1–12), and organizes existing hyperbolic transformers/LLMs, vision models, and multimodal models (Sec. 4), closing with challenges (Sec. 5). The paper's central empirical assertion, stated in the abstract and repeated in Sec. 4.1, is that recent hyperbolic models outperform Euclidean counterparts in LLM reasoning, VLM generalization, and cross-modal alignment while maintaining parameter efficiency.

Significance. If the central claim were supported by controlled comparisons, the survey would be a valuable roadmap for a promising direction: the technical catalogue of tangent-space, fully-hyperbolic, and refined operations (Sec. 3) is useful, and the taxonomy by modality and geometry mode (Table 1) is clear. The paper also honestly lists open problems, including the lack of large-scale hyperbolic pretraining and Riemannian training tools. However, the comparative evidence for geometric superiority is not yet established: the cited successes change multiple architectural components at once, and the strongest results come from the authors' own papers without benchmark tables or independent replication. The survey's value is therefore primarily as an organized review, not as a demonstration that hyperbolic geometry is the cause of the reported gains.

major comments (3)
  1. [Abstract; §4.1] The claim that hyperbolic geometry is a superior inductive bias is load-bearing, but the cited comparisons do not isolate geometry as the active ingredient. HELM (ref [47]) couples hyperbolic modules with a new Multi-Head Latent Attention and Mixture-of-Curvature-Experts; Hypformer (ref [115]) introduces hyperbolic linear attention plus readjustment/refinement blocks; LResNet (ref [50]) changes the residual connection; H-BLIP-2 (§4.3) changes the alignment loss and adds a stabilization term. In none of these cases does the Euclidean baseline receive the non-geometric counterpart of the new component with only the manifold removed. The reported gains are therefore equally compatible with the alternative explanation that the co-introduced architecture, not negative curvature, drives the improvement. The survey should either present controlled ablations or explicitly qualify the central claim.
  2. [§4.1–§4.3] The statement in §4.1 that 'HELM models outperformed Euclidean LLMs in billions-of-parameters scale on several benchmarks' cites only the authors' preprint [47] and is presented without any benchmark table, metric values, error bars, or training/evaluation budgets. Similar author-reported claims appear for H-BERT, Hypformer, HypLoRA, HyperCore/LViT/L-CLIP, and H-BLIP-2 in §§4.1–4.3. Since these results are the primary evidence for the survey's central claim, a survey should either include a consolidated results table with model sizes, datasets, and protocols, or flag these as unverified author-reported results and note the absence of independent replication.
  3. [§5] The paper's own assessment concedes that hyperbolic LLM pretraining is still at about one billion parameters, far below Euclidean-scale pretraining, that no hyperbolic vision foundation model has been pretrained, and that standard Riemannian tools such as a full AdamW equivalent and FlashAttention support are missing. These concessions directly undercut the abstract's 'while maintaining parameter efficiency' and the implied claim that hyperbolic geometry already enhances foundation models at scale. The abstract and conclusion should be reworded to present hyperbolic geometry as a promising but not yet established alternative, consistent with Section 5.
minor comments (4)
  1. [§3.1, Eq. (11)] Equation (11) contains a factor sqrt(K2/K2), which is identically 1; this is likely a typo for a curvature ratio involving K1 and K2, and should be corrected.
  2. [Throughout] There are several typos and inconsistent notations, including 'Poincaré aall model' (Appendix A), 'tengant' (Sec. 2.1), 'curvaure' (Sec. 3.2), 'Attnetion' (Sec. 3.1), and 'MEUR' (Sec. 4.3, should be MERU).
  3. [§5.4] The claim that 'The Hyp-GraphRAG model from HyperCore [49] provides a proof of concept' is not supported by any details or experimental result; either add a concrete reference to a result or remove the claim.
  4. [Figure 1] The caption says 'The red line shows the same geodesic under an isometry' but the two panels display different models of hyperbolic space; the caption should clarify what exactly is being compared.

Circularity Check

2 steps flagged · score 4.0 of 10

Survey's LLM/Transformer performance claims lean on self-citations (HELM, Hypformer), but independent works give the central thesis non-circular support.

  1. self citation load bearing [Section 4.1, Hyperbolic LLMs (paragraph on HELM, ref [47])]
    "When trained on the same datasets, HELM models outperformed Euclidean LLMs in billions-of-parameters scale on several benchmarks."

    The survey's central claim that hyperbolic geometry improves LLM capabilities at scale is supported in this passage solely by citation [47], HELM, whose author list includes the present survey's authors (He, Madhu, Yang, Ying). The survey does not reproduce or independently verify the benchmark comparison; the load-bearing evidence for the abstract's promise of 'improving LLMs' complex reasoning ability ... while maintaining parameter efficiency' is thus the authors' own prior work, not an external, machine-checked, or independently reproduced result.

  2. self citation load bearing [Section 4.1, Hyperbolic Transformers (paragraph on Hypformer, ref [115])]
    "Whether incorporating the GNN or not, Hypformer models outperformed Euclidean Transformer models on a variety of graph, image, and text tasks."

    Ref [115] (Hypformer) shares two authors with the present survey (Menglin Yang and Rex Ying). This sentence is the survey's primary evidence that fully hyperbolic Transformers beat Euclidean baselines across modalities, which is load-bearing for the abstract's claim that hyperbolic advances 'enhance foundation models'. No independent benchmark, external reproduction, or falsifiable analysis is cited in the survey for this specific claim. The circularity is partial because independent works (HAN, FNN, H-BERT) also appear in the same section.

full rationale

This paper is a literature survey rather than a derivation chain, so there is no equation-level circularity of the form where an output is identical to an input by construction. The mathematical background on low-distortion tree and power-law embeddings (refs [98,100]) is independent and non-circular. The central empirical thesis that hyperbolic geometry benefits foundation models is supported by many cited primary works, several of which (HELM [47], Hypformer [115], HyperCore [49], LResNet [50], HypLoRA [113]) overlap with the present authors. The two strongest scale-level claims for hyperbolic LLMs and fully hyperbolic Transformers rest on self-citations with no independent verification inside the survey, which is a load-bearing self-citation pattern rather than normal background citation. However, the survey also covers independent works (HAN, FNN, H-BERT, MERU, Hyp-ViT, H-BLIP-2, HyCoCLIP) that support the same general thesis, and Section 5 candidly concedes that hyperbolic pretraining remains far below Euclidean scale in model size and dataset size. The central claim therefore retains independent content and is not forced by the self-citation chain. The reviewer concern about confounded comparisons (multiple architectural components changing at once) is a legitimate experimental-design and correctness risk, but it is not circularity under the definitions used here.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

This survey introduces no new free parameters or invented entities. It relies on standard mathematical properties of hyperbolic space and on domain assumptions about data hierarchy and the reliability of reported experimental results, the latter of which is not independently verified.

assumptions (4)
  • standard math Hyperbolic space allows low-distortion embedding of trees and power-law distributions (e.g., Sarkar 2011; Sala et al. 2018).
    The survey uses this property to motivate hyperbolic models; it is a theorem from prior literature that the survey assumes.
  • domain assumption Real-world data in language, vision, and graphs is often hierarchical and scale-free.
    Stated in Section 1 with citations; if this is false, the motivation for hyperbolic geometry weakens substantially.
  • domain assumption The reported performance gains of hyperbolic models over Euclidean baselines are reliable and not due to confounding factors.
    The survey depends on these results for its central claim; they are taken from cited papers, many by the authors, without independent verification.
  • domain assumption Euclidean foundation models have fundamental limitations that are not addressable by better Euclidean training.
    This is the premise of the survey's motivation in Section 1; it is asserted rather than critically examined.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Hyperbolic Deep Learning for Foundation Models: A Survey." pith.science (2026). https://pith.science/paper/LF3VHLHE

@misc{pith2026250717787,
  author       = {Pith},
  title        = {Pith review of: Hyperbolic Deep Learning for Foundation Models: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LF3VHLHE}},
  note         = {Machine review of arXiv:2507.17787}
}
read the original abstract

Foundation models pre-trained on massive datasets, including large language models (LLMs), vision-language models (VLMs), and large multimodal models, have demonstrated remarkable success in diverse downstream tasks. However, recent studies have shown fundamental limitations of these models: (1) limited representational capacity, (2) lower adaptability, and (3) diminishing scalability. These shortcomings raise a critical question: is Euclidean geometry truly the optimal inductive bias for all foundation models, or could incorporating alternative geometric spaces enable models to better align with the intrinsic structure of real-world data and improve reasoning processes? Hyperbolic spaces, a class of non-Euclidean manifolds characterized by exponential volume growth with respect to distance, offer a mathematically grounded solution. These spaces enable low-distortion embeddings of hierarchical structures (e.g., trees, taxonomies) and power-law distributions with substantially fewer dimensions compared to Euclidean counterparts. Recent advances have leveraged these properties to enhance foundation models, including improving LLMs' complex reasoning ability, VLMs' zero-shot generalization, and cross-modal semantic alignment, while maintaining parameter efficiency. This paper provides a comprehensive review of hyperbolic neural networks and their recent development for foundation models. We further outline key challenges and research directions to advance the field.

Figures

Figures reproduced from arXiv: 2507.17787 by the authors.

Figure 1
Figure 1. Visualization of L 𝐾,𝑛 (top) and P 𝐾,𝑛 (bottom). The red line shows the same geodesic under an isometry. optimal performance when adapting to tasks that inherently exhibit hierarchical or power-law characteristics [17]. Finally, the scalability of foundation models poses increasing challenges due to their immense dimensionality and the explosive growth in computational [52]. Euclidean geometry exacerbates this issue… view at source ↗
Figure 3
Figure 3. Topics and words can be grouped in a dendrogram [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Embeddings of hyperbolic vision transformers push [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figures from the paper (1 more)
Figure 5
Figure 5. Figure 5: Visualization of hyperbolic entailment cone. This [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

136 extracted references · 48 canonical work pages

  1. [47]

    Neil He, Rishabh Anand, Hiren Madhu, Ali Maatouk, Smita Krishnaswamy, Lean- dros Tassiulas, Menglin Yang, and Rex Ying. 2025. HELM: Hyperbolic Large Lan- guage Models via Mixture-of-Curvature Experts. arXiv preprint arXiv:2505.24722 (2025)

  2. [115]

    Menglin Yang, Harshit Verma, Delvin Ce Zhang, Jiahong Liu, Irwin King, and Rex Ying. 2024. Hypformer: Exploring efficient transformer fully in hyperbolic space. In KDD. 3770–3781

  3. [50]

    Neil He, Menglin Yang, and Rex Ying. 2025. Lorentzian Residual Neural Net- works. In KDD

  4. [1]

    Gregorio Alanis-Lobato, Pablo Mier, and Miguel Andrade-Navarro. 2018. The latent geometry of the human protein interaction network. Bioinformatics 34, 16 (2018), 2826–2834

  5. [2]

    Muhammad Awais, Muzammal Naseer, Salman Khan, Rao Muhammad Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang, and Fahad Shahbaz Khan

  6. [3]

    Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer Normal- ization. arXiv arXiv:1607.06450 (2016)

  7. [4]

    Yushi Bai, Zhitao Ying, Hongyu Ren, and Jure Leskovec. 2021. Modeling hetero- geneous hierarchies with relation-specific hyperbolic cones. NeurIPS 34 (2021), 12316–12327

  8. [5]

    Federico Barbero, Alex Vitvitskyi, Christos Perivolaropoulos, Razvan Pascanu, and Petar Veličković. 2025. Round and Round We Go! What makes Rotary Positional Encodings useful? arXiv:2410.06205 (2025)

Show all 136 references
  1. [6]

    Ahmad Bdeir, Kristian Schwethelm, and Niels Landwehr. 2024. Fully Hyperbolic Convolutional Neural Networks for Computer Vision. In ICLR

  2. [7]

    Nithya Bhasker, Hattie Chung, Louis Boucherie, Vladislav Kim, Stefanie Speidel, and Melanie Weber. 2024. Contrastive poincaré maps for single-cell data analysis. In ICLR 2024 Workshop on Machine Learning for Genomics Explorations

  3. [8]

    Rishi Bommasani and et al. 2021. On the Opportunities and Risks of Foundation Models. arXiv preprint arXiv:2108.07258 (2021)

  4. [9]

    Joey Bose, Ariella Smofsky, Renjie Liao, Prakash Panangaden, and Will Hamilton

  5. [10]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. NeurIPS 33 (2020), 1877–1901

  6. [11]

    Davide Caffagni, Federico Cocchi, Luca Barsellotti, Nicholas Moratelli, Sara Sarto, Lorenzo Baraldi, Marcella Cornia, and Rita Cucchiara. 2024. The revolu- tion of multimodal large language models: a survey. arXiv:2402.12451 (2024)

  7. [12]

    Hanqun Cao, Cheng Tan, Zhangyang Gao, Yilun Xu, Guangyong Chen, Pheng- Ann Heng, and Stan Z Li. 2024. A survey on generative diffusion models. TKDE (2024)

  8. [13]

    Ines Chami, Zhitao Ying, Christopher Ré, and Jure Leskovec. 2019. Hyperbolic graph convolutional neural networks. In NeurIPS. 4868–4879

  9. [14]

    Jiayang Chen, Zhihang Hu, Siqi Sun, Qingxiong Tan, Yixuan Wang, Qinze Yu, Licheng Zong, Liang Hong, Jin Xiao, Tao Shen, et al. 2022. Interpretable RNA foundation model from unannotated data for highly accurate RNA structure and function predictions. arXiv:2204.00300 (2022)

  10. [15]

    Ming Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding, and Yaliang Li. 2020. Simple and deep graph convolutional networks. In ICML. PMLR, 1725–1735

  11. [16]

    Ricky T. Q. Chen and Yaron Lipman. 2024. Flow Matching on General Geometries. In ICLR

  12. [17]

    Weize Chen, Xu Han, Yankai Lin, Kaichen He, Ruobing Xie, Jie Zhou, and Zhiyuan Liu. 2024. Hyperbolic Pre-Trained Language Model. IEEE TASLP 32 (2024)

  13. [18]

    Weize Chen, Xu Han, Yankai Lin, Hexu Zhao, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2021. Fully Hyperbolic Neural Networks. arXiv:2105.14686 (2021)

  14. [19]

    Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. 2020. Improved base- lines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, (2020)

  15. [20]

    Yankai Chen, Menglin Yang, Yingxue Zhang, Mengchen Zhao, Ziqiao Meng, Jianye Hao, and Irwin King. 2022. Modeling Scale-free Graphs for Knowledge- aware Recommendation. WSDM (2022)

  16. [21]

    Jeffrey Cheng, Marc Marone, Orion Weller, Dawn Lawrie, Daniel Khashabi, and Benjamin Van Durme. 2024. Dated data: Tracing knowledge cutoffs in large language models. arXiv:2403.12958 (2024)

  17. [22]

    Andy Coenen, Emily Reif, Ann Yuan, Been Kim, Adam Pearce, Fernanda Viégas, and Martin Wattenberg. 2019. Visualizing and Measuring the Geometry of BERT. NuerIPS (2019)

  18. [23]

    Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah

  19. [24]

    Haotian Cui, Chloe Wang, Hassaan Maan, Kuan Pang, Fengning Luo, Nan Duan, and Bo Wang. 2024. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nature Methods 21, 8 (2024), 1470–1480

  20. [25]

    Jindou Dai, Yuwei Wu, Zhi Gao, and Yunde Jia. 2021. A Hyperbolic-to- Hyperbolic Graph Convolutional Network. arXiv:2104.06942 (2021), 154–163

  21. [26]

    Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

    Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. In NeurIPS

  22. [27]

    Valentin De Bortoli, Emile Mathieu, Michael Hutchinson, James Thornton, Yee Whye Teh, and Arnaud Doucet. 2022. Riemannian Score-Based Generative Modelling. In NeurIPS

  23. [28]

    DeepSeek-AI. 2024. DeepSeek-V3 Technical Report. arXiv:2412.19437 (2024)

  24. [29]

    Karan Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson, and Shan- mukha Ramakrishna Vedantam. 2023. Hyperbolic image-text representations. In ICML. PMLR, 7694–7731

  25. [30]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805 (2018)

  26. [31]

    Jiarui Ding and Aviv Regev. 2021. Deep generative model embedding of single- cell RNA-Seq profiles on hyperspheres and hyperbolic spaces. Nature communi- cations 12, 1 (2021), 2554

  27. [32]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al . 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11...

  28. [33]

    Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, and Steven Truitt. 2024. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv:2404.16130 (2024)

  29. [34]

    Aleksandr Ermolov, Leyla Mirvakhabova, Valentin Khrulkov, Nicu Sebe, and Ivan Oseledets. 2022. Hyperbolic vision transformers: Combining improvements in metric learning. In CVPR. 7409–7419

  30. [35]

    Jacob Fein-Ashley, Ethan Feng, and Minh Pham. 2024. HVT: A Comprehensive Vision Framework for Learning in Non-Euclidean Space. arXiv:2409.16897 (2024)

  31. [36]

    Xingcheng Fu, Yisen Gao, Yuecen Wei, Qingyun Sun, Hao Peng, Jianxin Li, and Xianxian Li. 2024. Hyperbolic Geometric Latent Diffusion Model for Graph Generation. ICML (2024)

  32. [37]

    Octavian Ganea, Gary Bécigneul, and Thomas Hofmann. 2018. Hyperbolic entailment cones for learning hierarchical embeddings. In ICML. PMLR, 1646– 1655

  33. [38]

    Octavian Ganea, Gary Bécigneul, and Thomas Hofmann. 2018. Hyperbolic neural networks. In NeurIPS. 5345–5355

  34. [39]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv:2312.10997 2 (2023)

  35. [40]

    Songwei Ge, Shlok Mishra, Simon Kornblith, Chun-Liang Li, and David Jacobs

  36. [41]

    Songwei Ge, Shlok Kumar Mishra, Simon Kornblith, Chun-Liang Li, and David Jacobs. 2022. Hyperbolic Contrastive Learning for Visual Representations be- yond Objects. ArXiv abs/2212.00653 (2022)

  37. [42]

    John Clifford Gower. 1985. Properties of Euclidean and non-Euclidean distance matrices. Linear algebra and its applications 67 (1985), 81–97

  38. [43]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The Llama 3 Herd of Models.arXiv:2407.21783 (2024)

  39. [44]

    Hyperbolic contrastive learning for visual representations beyond objects. In CVPR. 6840–6849

  40. [45]

    Zihao Guo, Qingyun Sun, Haonan Yuan, Xingcheng Fu, Min Zhou, Yisen Gao, and Jianxin Li. 2025. GraphMoRE: Mitigating Topological Heterogeneity via Mixture of Riemannian Experts. In AAAI

  41. [46]

    Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Mo- mentum contrast for unsupervised visual repre- sentation learning. In CVPR

  42. [48]

    Caglar Gulcehre, Misha Denil, Mateusz Malinowski, Ali Razavi, Razvan Pas- canu, Karl Moritz Hermann, Peter Battaglia, Victor Bapst, David Raposo, Adam Santoro, et al. 2019. Hyperbolic attention networks. In ICLR

  43. [49]

    Neil He, Menglin Yang, and Rex Ying. 2025. HyperCore: The Core Framework for Building Hyperbolic Foundation Models with Comprehensive Modules. arXiv:2504.08912 (2025)

  44. [51]

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising Diffusion Proba- bilistic Models. In NeurIPS

  45. [52]

    Neil He, Jiahong Liu, Buze Zhang, Ngoc Bui, Ali Maatouk, Menglin Yang, Irwin King, Melanie Weber, and Rex Ying. 2025. Position: Beyond Euclidean–Foundation Models Should Embrace Non-Euclidean Geometries. arXiv:2504.08896 (2025)

  46. [53]

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In ICLR

  47. [54]

    Chin-Wei Huang, Milad Aghajohari, Avishek Joey Bose, Prakash Panangaden, and Aaron Courville. 2022. Riemannian Diffusion Models. In NeurIPS

  48. [55]

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2025. KDD ’25, August 3–7, 2025, Toronto, ON, Canada Neil He, Hiren Madhu, Ngoc Bui, Menglin Yang, and Rex Ying A survey on hallucinatio...

  49. [56]

    Jordan Hoffmann and et al. 2022. Training compute-optimal large language models. In NeurIPS. Red Hook, NY, USA, Article 2176, 15 pages

  50. [57]

    Jaehyeong Jo and Hwang Sung Ju Lee, Seul and. 2022. Score-based Generative Modeling of Graphs via the System of Stochastic Differential Equations. In ICML

  51. [58]

    Ho, Percy Liang, and Arvind Narayanan

    Sayash Kapoor, Rishi Bommasani, Kevin Klyman, Shayne Longpre, Ashwin Ramaswami, Peter Cihon, Aspen Hopkins, Kevin Bankston, Stella Biderman, Miranda Bogen, Rumman Chowdhury, Alex Engler, Peter Henderson, Yacine Jer- nite, Seth Lazar, Stefano Maffulli, Alondra Nelson, Joelle Pi...

  52. [59]

    W Sean Kennedy, Onuttom Narayan, and Iraj Saniee. 2013. On the hyperbolicity of large-scale networks. arXiv:1307.0031 (2013)

  53. [60]

    Sergey Ioffe and Christian Szegedy. 2015. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. InICML. 448–456

  54. [61]

    Minseon Kim, Jihoon Tack, and Sung Ju Hwang. 2020. Adversarial Self- Supervised Contrastive Learning. In NeurIPS

  55. [62]

    Kingma and Max Welling

    Diederik P. Kingma and Max Welling. 2019. An Introduction to Variational Autoencoders. Foundations and Trends in Machine Learning, Vol. 12. 307–392 pages

  56. [63]

    Max Kochurov, Rasul Karimov, and Sergei Kozlukov. 2020. Geoopt: Riemannian Optimization in PyTorch. arXiv:2005.02819 (2020)

  57. [64]

    Bobak Kiani, Jason Wang, and Melanie Weber. 2024. Hardness of Learning Neural Networks under the Manifold Hypothesis. In NeurIPS

  58. [65]

    Marc Law, Renjie Liao, Jake Snell, and Richard Zemel. 2019. Lorentzian distance learning for hyperbolic representations. In ICML. PMLR, 3672–3681

  59. [66]

    Diego Lazcano, Nicolás Fredes, and Werner Creixell. 2021. Hyperbolic Genera- tive Adversarial Network. IEEE Access 9 (2021), 96309–96320

  60. [67]

    Matthew Le, Stephen Roller, Laetitia Papaxanthos, Douwe Kiela, and Maximilian Nickel. 2019. Inferring Concept Hierarchies from Text Corpora via Hyperbolic Embeddings. In ACL. 3231–3241

  61. [68]

    Dmitri Krioukov, Fragkiskos Papadopoulos, Maksim Kitsak, Amin Vahdat, and Marián Boguná. 2010. Hyperbolic geometry of complex networks. Physical Review E—Statistical, Nonlinear, and Soft Matter Physics 82, 3 (2010), 036106

  62. [69]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rock- täschel, et al. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In NeurIPS. 9459–9774

  63. [70]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Boot- strapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In ICML

  64. [71]

    Lingxiao Li, Yi Zhang, and Shuhui Wang. 2023. The Euclidean Space is Evil: Hyperbolic Attribute Editing for Few-shot Image Generation. In ICCV

  65. [72]

    Holden Lee, Jianfeng Lu, and Yixin Tan. 2022. Convergence for score-based generative modeling with polynomial complexity. In NeurIPS

  66. [73]

    Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. 2022. Flow Matching for Generative Modeling. arXiv:2210.02747 (2022)

  67. [74]

    Jiahong Liu, Menglin Yang, Min Zhou, Shanshan Feng, and Philippe Fournier- Viger. 2022. Enhancing Hyperbolic Graph Embeddings via Contrastive Learning. In NeurIPS 2nd SSL workshop

  68. [75]

    Qi Liu, Maximilian Nickel, and Douwe Kiela. 2019. Hyperbolic graph neural networks. In NeurIPS. 8230–8241

  69. [76]

    Nathan Linial, Eran London, and Yuri Rabinovich. 1995. The geometry of graphs and some of its algorithmic applications. Combinatorica 15, 2 (1995), 215–245

  70. [77]

    Aaron Lou, Isay Katsman, Qingxuan Jiang, Serge Belongie, Ser-Nam Lim, and Christopher De Sa. 2020. Differentiating through the Fréchet Mean. In ICML. 6393–6403

  71. [78]

    Qiyao Ma, Menglin Yang, Mingxuan Ju, Tong Zhao, Neil Shah, and Rex Ying

  72. [79]

    Paolo Mandica, Luca Franco, Konstantinos Kallidromitis, Suzanne Petryk, and Fabio Galasso. 2024. Hyperbolic Learning with Multimodal Large Language Models. arXiv:2408.05097 (2024)

  73. [80]

    Aaron Lou, Isay Katsman, Qingxuan Jiang, Serge Belongie, Ser-Nam Lim, and Christopher De Sa. 2020. Differentiating through the Fr \’echet Mean. arXiv:2003.00335 (2020)

  74. [81]

    Pascal Mettes, Mina Ghadimi Atigh, Martin Keller-Ressel, Jeffrey Gu, and Serena Yeung. 2024. Hyperbolic deep learning in computer vision: A survey. Interna- tional Journal of Computer Vision (2024), 1–25

  75. [82]

    Nina Miolane, Nicolas Guigui, Alice Le Brigant, Johan Mathe, Benjamin Hou, Yann Thanwerdas, Stefan Heyder, Olivier Peltre, Niklas Koep, Hadi Zaatiti, Hatem Hajri, Yann Cabanes, Thomas Gerald, Paul Chauchat, Christian Shew- make, Daniel Brooks, Bernhard Kainz, Claire Donnat, Su...

  76. [83]

    Devon Myers, Rami Mohawesh, Venkata Ishwarya Chellaboina, Anantha Lak- shmi Sathvik, Praveen Venkatesh, Yi-Hui Ho, Hanna Henshaw, Muna Al- hawawreh, David Berdik, and Yaser Jararweh. 2024. Foundation and large language models: fundamentals, challenges, opportunities, and socia...

  77. [84]

    Yoshihiro Nagano, Shoichiro Yamaguchi, Yasuhiro Fujita, and Masanori Koyama

  78. [85]

    Maddison, Ryota Tomioka, and Yee Whye Teh

    Emile Mathieu, Charline Le Lan, Chris J. Maddison, Ryota Tomioka, and Yee Whye Teh. 2019. Continuous Hierarchical Representations with Poincaré Variational Auto-Encoders. In NeurIPS

  79. [86]

    Maximillian Nickel and Douwe Kiela. 2018. Learning Continuous Hierarchies in the Lorentz Model of Hyperbolic Geometry. In ICML. 3779–3788

  80. [87]

    Keiron O’Shea and Ryan Nash. 2015. An Introduction to Convolutional Neural Networks. arXiv:1511.08458 (2015)

  81. [88]

    Avik Pal, Max van Spengler, Guido Maria D’Amely di Melendugno, Alessandro Flaborea, Fabio Galasso, and Pascal Mettes. 2025. Compositional Entailment Learning for Hyperbolic Vision-Language Models. ICLR (2025)

  82. [89]

    Fragkiskos Papadopoulos, Dmitri Krioukov, Marián Boguná, and Amin Vah- dat. 2010. Greedy forwarding in dynamic scale-free networks embedded in hyperbolic metric spaces. In 2010 Proceedings IEEE Infocom . IEEE, 1–9

  83. [90]

    Wei Peng, Tuomas Varanka, Abdelrahman Mostafa, Henglin Shi, and Guoying Zhao. 2021. Hyperbolic deep neural networks: A survey. TPAMI (2021)

  84. [91]

    Maximillian Nickel and Douwe Kiela. 2017. Poincaré embeddings for learning hierarchical representations. In NeurIPS. 6338–6347

  85. [92]

    Aleksandar Poleksic. 2023. Hyperbolic matrix factorization improves prediction of drug-target associations. Scientific Reports 13, 1 (2023), 959

  86. [94]

    Eric Qu and Dongmian Zou. 2022. Lorentzian fully hyperbolic generative adversarial network. arXiv:2201.12825 (2022)

  87. [95]

    Radford, Kim Jong Wook Alec, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. arXiv:2103.000...

  88. [96]

    Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deep- Speed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters. In KDD. 3505–3506

  89. [97]

    Xavier Pennec. 2006. Intrinsic Statistics on Riemannian Manifolds: Basic Tools for Geometric Measurements. JMIV 25, 1 (2006), 127–154

  90. [98]

    Frederic Sala, Chris De Sa, Albert Gu, and Christopher Re. 2018. Representation Tradeoffs for Hyperbolic Embeddings. In ICML. 4460–4469

  91. [99]

    Lillemark, Christian Shewmake, Abby Bertics, Xavier Pennec, and Nina Miolane

    Sophia Sanborn, Johan Mathe, Mathilde Papillon, Domas Buracas, Hansen J. Lillemark, Christian Shewmake, Abby Bertics, Xavier Pennec, and Nina Miolane

  92. [100]

    Rik Sarkar. 2011. Low distortion delaunay embedding of trees in hyperbolic plane. In International Symposium on Graph Drawing . Springer, 355–366

  93. [101]

    Ryohei Shimizu, Yusuke Mukuta, and Tatsuya Harada. 2020. Hyperbolic Neural Networks++. In ICLR

  94. [102]

    Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Ste- fano Ermon, and Ben Poole. 2021. Score-Based Generative Modeling through Stochastic Differential Equations. In ICLR

  95. [103]

    Salem Said, Lionel Bombrun, and Yannick Berthoumieu. 2014. New Riemannian Priors on the Univariate Normal Model. Entropy 16, 7 (2014), 4015–4031

  96. [104]

    Xunzhu Tang, Saad Ezzini, Haoye Tian, Yewei Song, Jacques Klein, Tegawende F Bissyande, et al. 2023. Hyperbolic code retrieval: a novel approach for efficient code search using hyperbolic space embeddings. arXiv:2308.15234 (2023)

  97. [105]

    Alexandru Tifrea, Gary Bécigneul, and Octavian-Eugen Ganea. 2018. Poincaré glove: Hyperbolic word embeddings. arXiv:1810.06546 (2018)

  98. [106]

    arXiv:2407.09468 (2024)

    Beyond Euclid: An Illustrated Guide to Modern Machine Learning with Geometric, Topological, and Algebraic Structures. arXiv:2407.09468 (2024)

  99. [107]

    Max van Spengler, Erwin Berkhout, and Pascal Mettes. 2023. Poincaré ResNet. CVPR (2023)

  100. [108]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NeurIPS. 5998–6008

  101. [109]

    Amelia Villegas-Morcillo, Stavros Makrodimitris, Roeland C H J van Ham, Angel M Gomez, Victoria Sanchez, and Marcel J T Reinders. 2020. Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function. Bioinformatics 37, ...

  102. [110]

    Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2021. RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv:2104.09864 (2021)

  103. [111]

    Lingfeng Wen, Xuan Tang, Mingjie Ouyang, Xiangxiang Shen, Jian Yang, Daxin Zhu, Mingsong Chen, and Xian Wei. 2024. Hyperbolic Graph Diffusion Model. In AAAI. Hyperbolic Deep Learning for Foundation Models: A Survey KDD ’25, August 3–7, 2025, Toronto, ON, Canada

  104. [112]

    Jiahao Xie, Xiaohang Zhan, Ziwei Liu, Yew Soon Ong, and Chen Change Loy

  105. [113]

    Abraham Albert Ungar. 2008. A gyrovector space approach to hyperbolic geometry. Synthesis Lectures on Mathematics and Statistics 1, 1 (2008), 1–194

  106. [114]

    Menglin Yang, Zhihao Li, Min Zhou, Jiahong Liu, and Irwin King. 2022. Hicf: Hyperbolic informative collaborative filtering. In KDD. 2212–2221

  107. [116]

    Menglin Yang, Min Zhou, Marcus Kalander, Zengfeng Huang, and Irwin King

  108. [117]

    Xinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong, and Lei Li. 2021. Dense contrastive learning for self-supervised visual pre-training. In CVPR

  109. [118]

    Menglin Yang, Min Zhou, Jiahong Liu, Defu Lian, and Irwin King. 2022. HRCF: Enhancing Collaborative Filtering via Hyperbolic Geometric Regularization. In WebConf

  110. [119]

    Menglin Yang, Min Zhou, Hui Xiong, and Irwin King. 2022. Hyperbolic Temporal Network Embedding. TKDE (2022)

  111. [120]

    Yamada, and Ziming Zhang

    Yun Yue, Fangzhou Lin, Kazunori D. Yamada, and Ziming Zhang. 2023. Hyper- bolic Contrastive Learning. arXiv:2302.01409 (2023)

  112. [121]

    Menglin Yang, Aosong Feng, Bo Xiong, Jihong Liu, Irwin King, and Rex Ying

  113. [122]

    ICML LLM Cognition Workshop (2024)

    Hyperbolic Fine-tuning for Large Language Models. ICML LLM Cognition Workshop (2024)

  114. [123]

    Yiding Zhang, Xiao Wang, Chuan Shi, Nian Liu, and Guojie Song. 2021. Lorentzian Graph Convolutional Networks. In WebConf. 1249–1261

  115. [124]

    Yifei Zhang, Hao Zhu, Menglin Yang, Jiahong Liu, Rex Ying, Irwin King, and Piotr Koniusz. 2025. Understanding and mitigating hyperbolic dimensional collapse in graph contrastive learning. In KDD. 1984–1995

  116. [125]

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv:2303.18223 1, 2 (2023)

  117. [126]

    Discrete-time Temporal Network Embedding via Implicit Hierarchical Learning in Hyperbolic Space. In KDD. 1975–1985

  118. [127]

    Menglin Yang, Min Zhou, Zhihao Li, Jiahong Liu, Lujia Pan, Hui Xiong, and Irwin King. 2022. Hyperbolic graph neural networks: a review of methods and applications. arXiv:2202.13852 (2022)

  119. [131]

    Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu. 2024. Vision-Language Models for Vision Tasks: A Survey. IEEE TPAMI 46, 8 (2024), 5625–5644

  120. [132]

    Yiding Zhang, Xiao Wang, Chuan Shi, Xunqiang Jiang, and Yanfang Fanny Ye

  121. [133]

    TBD (2021)

    Hyperbolic graph attention network. TBD (2021)

  122. [137]

    Shichao Zhu, Shirui Pan, Chuan Zhou, Jia Wu, Yanan Cao, and Bin Wang. 2020. Graph Geometry Interaction Learning. In NeurIPS, Vol. 33. 7548–7558. A Hyperbolic Geometry and Additional Works A.1 Hyperbolic Geometry Lorentz model. An𝑛-dimensional Lorentz model is a Riemann- ian ma...

  123. [2019]

    A wrapped normal distribution on hyperbolic space for gradient-based learning. In ICML. PMLR, 4693–4702

  124. [2020]

    Latent variable modelling with hyperbolic normalizing flows. In ICML. PMLR, 1045–1055

  125. [2021]

    In NeurIPS

    Unsupervised object-level representation learning from scene images. In NeurIPS

  126. [2023]

    TPAMI (2023)

    Diffusion models in vision: A survey. TPAMI (2023)

  127. [2024]

    arXiv:2411.13865 (2024)

    HARec: Hyperbolic graph-llm alignment for exploration and exploitation in recommender systems. arXiv:2411.13865 (2024)

  128. [2025]

    TPAMI (2025)

    Foundation Models Defining a New Era in Vision: a Survey and Outlook. TPAMI (2025)

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.