Pith. sign in

REVIEW 2 major objections 73 references

Stable current-carrying rings called kinky vortons exist in the Z2-symmetric two-Higgs-doublet model and match thin-string predictions.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 20:28 UTC pith:DA3FMY7H

load-bearing objection Abstract promises stable kinky vortons in Z2 2HDM with thin-string agreement, but the supplied full text is the wrong paper (UNCHA/cs.CV), so nothing technical can be checked. the 2 major comments →

arxiv 2603.22040 v2 pith:DA3FMY7H submitted 2026-03-23 hep-ph astro-ph.COhep-th

Kinky vortons in the 2HDM

classification hep-ph astro-ph.COhep-th PACS 11.27.+d12.60.Fr98.80.Cq
keywords kinky vortonstwo-Higgs-doublet modelZ2 symmetrycurrent-carrying stringsthin-string approximationelastic string formalismdomain wallssolitons
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper constructs and studies two-dimensional current-carrying ring solutions, known as kinky vortons, inside the Z2-symmetric global two-Higgs-doublet model. It shows that several of these rings are dynamically stable and survive even when they are kicked with non-axisymmetric perturbations. Their equilibrium sizes and the frequencies at which they oscillate are accurately reproduced by the thin-string approximation together with the elastic-string formalism. Because the model is a simplified but still phenomenologically motivated extension of the Standard Model, the existence of these stable objects suggests that vortons can appear in more realistic Higgs sectors as well. The work also finds a composite domain-wall arrangement in which condensates sit on secondary walls living on a primary Z2 wall, offering a possible route by which similar ring-like defects could form in three dimensions.

Core claim

Multiple dynamically stable kinky vortons exist in the Z2-symmetric global 2HDM; they remain intact under non-axially symmetric perturbations, and both their equilibrium radii and their dynamical oscillation frequencies are correctly captured by the thin-string approximation and the elastic-string formalism.

What carries the argument

Kinky vortons—two-dimensional, current-carrying ring solutions of the Z2-symmetric global two-Higgs-doublet model—together with the thin-string approximation and elastic-string formalism that predict their radii and frequencies.

Load-bearing premise

That the simplified two-dimensional Z2-symmetric global model is a good enough proxy for the more realistic U(1)-symmetric and ultimately three-dimensional gauged 2HDM that the stability found here will carry over.

What would settle it

A numerical evolution of the full three-dimensional U(1)-symmetric 2HDM that either fails to produce any long-lived current-carrying rings or produces rings whose radii and oscillation frequencies deviate sharply from the thin-string and elastic-string predictions.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Vorton solutions become viable objects inside a phenomenologically motivated extension of the Standard Model.
  • The Z2-symmetric theory can serve as a computationally cheaper laboratory for studying vortons that should also exist in the U(1)-symmetric 2HDM.
  • A composite domain-wall configuration supplies a concrete mechanism by which kinky-vorton-like defects could form in three spatial dimensions.
  • Thin-string and elastic-string analytic tools can be trusted for quantitative predictions of equilibrium size and oscillation spectrum in this class of models.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the same thin-string agreement persists once the U(1) symmetry is restored and the theory is gauged, kinky vortons could leave observable cosmological or collider signatures in two-Higgs-doublet scenarios.
  • The secondary-wall condensate structure may provide a template for constructing higher-dimensional solitons that carry both topological charge and persistent current.
  • Stability under non-axisymmetric perturbations suggests that these rings are not fine-tuned artefacts of axial symmetry and could therefore appear in a thermal bath or after a phase transition.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 0 minor

Summary. The manuscript claims to construct and analyse two-dimensional current-carrying ring solutions (kinky vortons) in the Z2-symmetric global two-Higgs-doublet model. It asserts the existence of multiple dynamically stable configurations that survive non-axially symmetric perturbations, that their equilibrium radii and oscillation frequencies are accurately captured by the thin-string approximation and elastic-string formalism, and that these solutions establish the viability of vortons in a phenomenologically motivated SM extension while serving as a tractable proxy for the U(1)-symmetric 2HDM. A composite domain-wall configuration with localized condensates on secondary walls is also reported as a possible three-dimensional formation mechanism.

Significance. If the claimed constructions, stability tests and quantitative thin-string/elastic-string agreement hold, the work would supply the first concrete evidence that current-carrying vortons exist and remain dynamically stable inside a realistic multi-Higgs extension of the Standard Model. That would be a non-trivial advance for both topological-defect cosmology and 2HDM phenomenology, and the proposed Z2 proxy could lower the computational barrier to studying the more expensive U(1) case. The composite domain-wall observation would further suggest a concrete pathway from domain walls to vorton-like objects in three dimensions. These strengths cannot be verified from the material supplied.

major comments (2)
  1. The full manuscript body supplied under this arXiv identifier is an unrelated computer-vision paper (UNCHA / arXiv:2603.22042). No field equations, energy functionals, numerical profiles, stability spectra, thin-string comparisons or composite-domain-wall constructions for the Z2-symmetric 2HDM appear. Consequently the central claims of existence, dynamical stability under non-axial perturbations, and quantitative agreement of radii and frequencies with the thin-string and elastic-string formalisms cannot be checked.
  2. The abstract's key proxy claim—that results in the two-dimensional Z2-symmetric global theory transfer to the phenomenologically relevant U(1)-symmetric (and ultimately gauged, three-dimensional) 2HDM—is left unsubstantiated because the supporting calculations are absent from the supplied text. Without those sections the load-bearing assumption of the paper remains untested.

Circularity Check

0 steps flagged

No circularity can be assessed or found: supplied full text is the unrelated UNCHA/cs.CV paper, not the kinky-vortons hep-ph manuscript; abstract alone shows no self-definitional or fitted-prediction loop.

full rationale

The CACHEABLE PAPER SOURCE CONTEXT and FULL MANUSCRIPT TEXT blocks contain the complete text of an entirely different work (UNCHA, arXiv:2603.22042, hyperbolic VLMs). No equations, thin-string approximations, elastic-string frequencies, stability tests under non-axial perturbations, or composite domain-wall constructions from the claimed 2HDM kinky-vorton paper appear. Consequently the derivation chain of the target paper cannot be walked. From the abstract alone the strongest claims (existence of multiple dynamically stable current-carrying rings whose radii and frequencies match the thin-string/elastic-string predictions) are presented as numerical constructions and external checks, not as tautological redefinitions of fitted inputs or self-citation uniqueness theorems. No self-definitional step, fitted-input-called-prediction, or load-bearing self-citation is quotable. Score 0 is therefore required; residual proxy-assumption risk is outside the circularity criterion.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 2 invented entities

Abstract-only review. Load-bearing background includes standard classical field theory of global defects, the existence of the Z2-symmetric global 2HDM potential, and the validity of thin-string and elastic-string reductions. No free parameters or invented particles can be extracted beyond the model choice itself. The main modeling leap is treating the 2D Z2 theory as a proxy for more realistic 2HDM vortons.

axioms (3)
  • domain assumption The Z2-symmetric global two-Higgs-doublet model admits topological kink and current-carrying string-like solutions that can close into rings.
    Implicit foundation of the entire construction; stated as the setting of the paper in the abstract.
  • domain assumption Thin-string approximation and elastic-string formalism accurately describe equilibrium radii and oscillation frequencies of these rings.
    Used as the analytic benchmark that the numerical solutions are said to match; standard in cosmic-string literature but assumed applicable here.
  • ad hoc to paper Stability under non-axially symmetric perturbations in 2D is informative for the viability of vortons in the related U(1)-symmetric and 3D settings.
    The proxy claim in the abstract rests on this transferability assumption, which is not demonstrated in the abstract.
invented entities (2)
  • kinky vortons in the Z2-symmetric global 2HDM no independent evidence
    purpose: Provide concrete, stable current-carrying ring solutions inside a phenomenologically motivated Higgs sector and a cheaper proxy for U(1) 2HDM vortons.
    Named solutions constructed in this work; independent evidence would require published field profiles, spectra, or cosmological signatures outside this abstract.
  • composite domain wall with localized condensates on secondary walls no independent evidence
    purpose: Suggest a 3D formation pathway for kinky-vorton-like defects.
    Additional configuration identified in the abstract; no external handle given here.

pith-pipeline@v1.1.0-grok45 · 28782 in / 2593 out tokens · 29126 ms · 2026-07-13T20:28:20.096206+00:00 · methodology

0 comments
read the original abstract

We construct and analyse two-dimensional, current-carrying ring solutions, known as kinky vortons, in the $\mathbb{Z}_2$-symmetric global two-Higgs-doublet model (2HDM). We demonstrate the existence of multiple dynamically stable configurations that persist under non-axially symmetric perturbations. These solutions are described with high accuracy by the thin string approximation and elastic string formalism, which correctly capture both their equilibrium radii and dynamical oscillation frequencies. Kinky vortons in the $\mathbb{Z}_2$-symmetric theory establish the viability of vorton solutions in a phenomenologically motivated extension of the Standard Model, and should provide a computationally tractable proxy for vortons in the $U(1)$-symmetric 2HDM. In addition, we identify a composite domain wall configuration in which localized condensates are supported on secondary domain walls existing on a $\mathbb{Z}_2$ wall, suggesting a mechanism by which kinky-vorton-like defects could arise in a three dimensional setting.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

73 extracted references · 8 linked inside Pith

  1. [1]

    Clip under the microscope: A fine-grained analysis of multi-object representation

    Reza Abbasi, Ali Nazari, Aminreza Sefid, Moham- madali Banayeeanzade, Mohammad Hossein Rohban, and Mahdieh Soleymani Baghshah. Clip under the microscope: A fine-grained analysis of multi-object representation. In CVPR, 2025. 1, 2, 7, 11

  2. [2]

    Hyperbolic image segmentation

    Mina Ghadimi Atigh, Julian Schoep, Erman Acar, Nanne van Noord, and Pascal Mettes. Hyperbolic image segmentation. InCVPR, 2022. 2, 3, 4

  3. [3]

    Springer Science & Business Media,

    Martin R Bridson and Andr ´e Haefliger.Metric spaces of non-positive curvature. Springer Science & Business Media,

  4. [4]

    Coyo-700m: Image-text pair dataset.https : / / github

    Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. Coyo-700m: Image-text pair dataset.https : / / github . com / kakaobrain/coyo-dataset, 2022. Dataset. 10

  5. [5]

    Hyperbolic graph convolutional neural networks

    Ines Chami, Zhitao Ying, Christopher R ´e, and Jure Leskovec. Hyperbolic graph convolutional neural networks. NeurIPS, 32, 2019. 3

  6. [6]

    From trees to continuous embeddings and back: Hyperbolic hierarchical clustering.NeurIPS, 2020

    Ines Chami, Albert Gu, Vaggos Chatziafratis, and Christo- pher R ´e. From trees to continuous embeddings and back: Hyperbolic hierarchical clustering.NeurIPS, 2020. 1

  7. [7]

    Horopca: Hyperbolic dimensionality reduction via horo- spherical projections

    Ines Chami, Albert Gu, Dat P Nguyen, and Christopher R´e. Horopca: Hyperbolic dimensionality reduction via horo- spherical projections. InInternational Conference on Ma- chine Learning, pages 1419–1429. PMLR, 2021. 13, 15

  8. [8]

    Towards understand- ing hierarchical learning: Benefits of neural representations

    Minshuo Chen, Yu Bai, Jason D Lee, Tuo Zhao, Huan Wang, Caiming Xiong, and Richard Socher. Towards understand- ing hierarchical learning: Benefits of neural representations. NeurIPS, 2020. 1, 10

  9. [9]

    Imagenet: A large-scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. InCVPR, 2009. 7, 11, 16

  10. [10]

    Hyper- bolic image-text representations

    Karan Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson, and Shanmukha Ramakrishna Vedantam. Hyper- bolic image-text representations. InICML, 2023. 1, 2, 3, 4, 5, 6, 7, 8, 10, 14, 15

  11. [11]

    Embedding text in hyperbolic spaces

    Bhuwan Dhingra, Christopher Shallue, Mohammad Norouzi, Andrew Dai, and George Dahl. Embedding text in hyperbolic spaces. InProceedings of the Twelfth Workshop on Graph-Based Methods for Natural Language Processing (TextGraphs-12), 2018. 1, 3

  12. [12]

    An image is worth 16x16 words: Trans- formers for image recognition at scale.ICLR, 2021

    Alexey Dosovitskiy. An image is worth 16x16 words: Trans- formers for image recognition at scale.ICLR, 2021. 10

  13. [13]

    Write a classifier: Zero-shot learning using purely textual descriptions

    Mohamed Elhoseiny, Babak Saleh, and Ahmed Elgammal. Write a classifier: Zero-shot learning using purely textual descriptions. InCVPR, 2013. 10

  14. [14]

    The pascal visual object classes (voc) challenge.IJCV, 2010

    Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge.IJCV, 2010. 7, 11, 17

  15. [15]

    Hyperbolic active learning for semantic segmen- tation under domain shift.ICML, 2023

    Luca Franco, Paolo Mandica, Konstantinos Kallidromitis, Devin Guillory, Yu-Teng Li, Trevor Darrell, and Fabio Galasso. Hyperbolic active learning for semantic segmen- tation under domain shift.ICML, 2023. 2, 3, 4

  16. [16]

    Hyperbolic entailment cones for learning hierarchical em- beddings

    Octavian Ganea, Gary B ´ecigneul, and Thomas Hofmann. Hyperbolic entailment cones for learning hierarchical em- beddings. InICML, 2018. 1, 3

  17. [17]

    Improving zero-shot gen- eralization and robustness of multi-modal models

    Yunhao Ge, Jie Ren, Andrew Gallagher, Yuxiao Wang, Ming-Hsuan Yang, Hartwig Adam, Laurent Itti, Balaji Lak- shminarayanan, and Jiaping Zhao. Improving zero-shot gen- eralization and robustness of multi-modal models. InCVPR,

  18. [18]

    Semi-supervised learning by entropy minimization.NeurIPS, 2004

    Yves Grandvalet and Yoshua Bengio. Semi-supervised learning by entropy minimization.NeurIPS, 2004. 5

  19. [19]

    Lvis: A dataset for large vocabulary instance segmentation

    Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. InCVPR,

  20. [20]

    Pay attention to your neighbours: Training-free open-vocabulary semantic segmentation

    Sina Hajimiri, Ismail Ben Ayed, and Jose Dolz. Pay attention to your neighbours: Training-free open-vocabulary semantic segmentation. InWACV, 2025. 13

  21. [21]

    Position: Beyond euclidean–foundation models should embrace non-euclidean geometries.arXiv preprint arXiv:2504.08896, 2025

    Neil He, Jiahong Liu, Buze Zhang, Ngoc Bui, Ali Maa- touk, Menglin Yang, Irwin King, Melanie Weber, and Rex Ying. Position: Beyond euclidean–foundation models should embrace non-euclidean geometries.arXiv preprint arXiv:2504.08896, 2025. 1

  22. [22]

    Hyperbolic deep learning for foundation models: A survey

    Neil He, Hiren Madhu, Ngoc Bui, Menglin Yang, and Rex Ying. Hyperbolic deep learning for foundation models: A survey. InProceedings of the 31st ACM SIGKDD Confer- ence on Knowledge Discovery and Data Mining V . 2, 2025. 3

  23. [23]

    Hypercore: The core framework for building hyperbolic foundation models with comprehensive modules.arXiv preprint arXiv:2504.08912,

    Neil He, Menglin Yang, and Rex Ying. Hypercore: The core framework for building hyperbolic foundation models with comprehensive modules.arXiv preprint arXiv:2504.08912,

  24. [24]

    Fine-grained image classifi- cation via combining vision and language

    Xiangteng He and Yuxin Peng. Fine-grained image classifi- cation via combining vision and language. InCVPR, 2017. 2

  25. [25]

    Some demonstrations of the effects of structural descriptions in mental imagery.Cognitive Science,

    Geoffrey Hinton. Some demonstrations of the effects of structural descriptions in mental imagery.Cognitive Science,

  26. [26]

    How to represent part-whole hierarchies in a neural network.Neural Computation, 2023

    Geoffrey Hinton. How to represent part-whole hierarchies in a neural network.Neural Computation, 2023. 1

  27. [27]

    Pixel-bert: Aligning image pixels with text by deep multi-modal transformers.arXiv preprint arXiv:2004.00849, 2020

    Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. Pixel-bert: Aligning image pixels with text by deep multi-modal transformers.arXiv preprint arXiv:2004.00849, 2020. 2

  28. [28]

    Intriguing properties of hyperbolic embeddings in vision-language models.TMLR,

    Sarah Ibrahimi, Mina Ghadimi Atigh, Nanne Van Noord, Pascal Mettes, and Marcel Worring. Intriguing properties of hyperbolic embeddings in vision-language models.TMLR,

  29. [29]

    Open- clip, 2021

    Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Han- naneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. Open- clip, 2021. 10

  30. [30]

    Compositional generaliza- tion through abstract representations in human and artificial neural networks.NeurIPS, 2022

    Takuya Ito, Tim Klinger, Doug Schultz, John Murray, Michael Cole, and Mattia Rigotti. Compositional generaliza- tion through abstract representations in human and artificial neural networks.NeurIPS, 2022. 1

  31. [31]

    Scaling up visual and vision-language representation learning with noisy text supervision

    Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. InICML, pages 4904–

  32. [32]

    Refclip: A universal teacher for weakly supervised referring expression comprehension

    Lei Jin, Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Guannan Jiang, Annan Shu, and Rongrong Ji. Refclip: A universal teacher for weakly supervised referring expression comprehension. InCVPR, 2023. 2

  33. [33]

    Fineclip: Self- distilled region-based clip for better fine-grained understand- ing.NeurIPS, 2024

    Dong Jing, Xiaolong He, Yutian Luo, Nanyi Fei, Wei Wei, Huiwen Zhao, Zhiwu Lu, et al. Fineclip: Self- distilled region-based clip for better fine-grained understand- ing.NeurIPS, 2024. 13

  34. [34]

    Deep visual-semantic align- ments for generating image descriptions

    Andrej Karpathy and Li Fei-Fei. Deep visual-semantic align- ments for generating image descriptions. InCVPR, 2015. 6, 10

  35. [35]

    Hyperbolic im- age embeddings

    Valentin Khrulkov, Leyla Mirvakhabova, Evgeniya Usti- nova, Ivan Oseledets, and Victor Lempitsky. Hyperbolic im- age embeddings. InCVPR, 2020. 1, 3

  36. [36]

    Unifying visual-semantic embeddings with multimodal neu- ral language models.arXiv preprint arXiv:1411.2539, 2014

    Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel. Unifying visual-semantic embeddings with multimodal neu- ral language models.arXiv preprint arXiv:1411.2539, 2014. 2

  37. [37]

    The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.IJCV,

    Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Ui- jlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al. The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.IJCV,

  38. [38]

    Inferring concept hierarchies from text corpora via hyperbolic embeddings

    Matthew Le, Stephen Roller, Laetitia Papaxanthos, Douwe Kiela, and Maximilian Nickel. Inferring concept hierarchies from text corpora via hyperbolic embeddings. InProceed- ings of the 57th annual meeting of the association for com- putational linguistics, 2019. 3, 5

  39. [39]

    Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in neural infor- mation processing systems, 34:9694–9705, 2021

    Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in neural infor- mation processing systems, 34:9694–9705, 2021. 1, 2

  40. [40]

    Microsoft coco: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 6, 7, 10, 11, 13, 15, 17, 22

  41. [41]

    Hyperbolic graph neural networks.NeurIPS, 32, 2019

    Qi Liu, Maximilian Nickel, and Douwe Kiela. Hyperbolic graph neural networks.NeurIPS, 32, 2019. 3

  42. [42]

    Sgdr: Stochastic gradient descent with warm restarts.ICLR, 2017

    Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts.ICLR, 2017. 10

  43. [43]

    Decoupled weight decay regularization.ICLR, 2019

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization.ICLR, 2019. 10

  44. [44]

    Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.NeurIPS, 32, 2019

    Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.NeurIPS, 32, 2019. 2

  45. [45]

    Rec- tifier nonlinearities improve neural network acoustic models

    Andrew L Maas, Awni Y Hannun, Andrew Y Ng, et al. Rec- tifier nonlinearities improve neural network acoustic models. InICML, 2013. 5

  46. [46]

    Hyperbolic learning with multimodal large language models

    Paolo Mandica, Luca Franco, Konstantinos Kallidromitis, Suzanne Petryk, and Fabio Galasso. Hyperbolic learning with multimodal large language models. InECCV, 2024. 2, 3, 4

  47. [47]

    Wordnet: a lexical database for english

    George A Miller. Wordnet: a lexical database for english. Communications of the ACM, 1995. 11

  48. [48]

    Poincar ´e embeddings for learning hierarchical representations.NeurIPS, 2017

    Maximillian Nickel and Douwe Kiela. Poincar ´e embeddings for learning hierarchical representations.NeurIPS, 2017. 1, 2

  49. [49]

    Compositional entailment learning for hyperbolic vision-language models.ICLR, 2024

    Avik Pal, Max van Spengler, Guido Maria D’Amely di Me- lendugno, Alessandro Flaborea, Fabio Galasso, and Pascal Mettes. Compositional entailment learning for hyperbolic vision-language models.ICLR, 2024. 1, 2, 3, 4, 5, 6, 7, 8, 10, 11, 14, 15, 21

  50. [50]

    Hyperbolic deep neural networks: A survey.IEEE Transactions on pattern analysis and machine intelligence, 2021

    Wei Peng, Tuomas Varanka, Abdelrahman Mostafa, Henglin Shi, and Guoying Zhao. Hyperbolic deep neural networks: A survey.IEEE Transactions on pattern analysis and machine intelligence, 2021. 2

  51. [51]

    Kosmos-2: Ground- ing multimodal large language models to the world.arXiv preprint arXiv:2306.14824, 2023

    Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. Kosmos-2: Ground- ing multimodal large language models to the world.arXiv preprint arXiv:2306.14824, 2023. 6, 7, 10

  52. [52]

    What does a platypus look like? generating customized prompts for zero-shot image classification

    Sarah Pratt, Ian Covert, Rosanne Liu, and Ali Farhadi. What does a platypus look like? generating customized prompts for zero-shot image classification. InCVPR, 2023. 2

  53. [53]

    Learn- ing transferable visual models from natural language super- vision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. InICML, 2021. 1, 2, 6, 7, 8, 10, 14, 15

  54. [54]

    Accept the modality gap: An exploration in the hyperbolic space

    Sameera Ramasinghe, Violetta Shevchenko, Gil Avraham, and Ajanthan Thalaiyasingam. Accept the modality gap: An exploration in the hyperbolic space. InCVPR, 2024. 1, 2, 3, 5, 6, 7, 8, 10, 14, 15

  55. [55]

    Joint image-text representation by gaussian visual-semantic embedding

    Zhou Ren, Hailin Jin, Zhe Lin, Chen Fang, and Alan Yuille. Joint image-text representation by gaussian visual-semantic embedding. InProceedings of the 24th ACM international conference on Multimedia, 2016. 2

  56. [56]

    Imagenet large scale visual recognition challenge.IJCV, 2015

    Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, San- jeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge.IJCV, 2015. 6, 8, 11, 12, 16

  57. [57]

    Clip for all things zero-shot sketch-based image retrieval, fine- grained or not

    Aneeshan Sain, Ayan Kumar Bhunia, Pinaki Nath Chowd- hury, Subhadeep Koley, Tao Xiang, and Yi-Zhe Song. Clip for all things zero-shot sketch-based image retrieval, fine- grained or not. InCVPR, 2023. 2

  58. [58]

    Representation tradeoffs for hyperbolic embeddings

    Frederic Sala, Chris De Sa, Albert Gu, and Christopher R´e. Representation tradeoffs for hyperbolic embeddings. In ICML, 2018. 1, 3

  59. [59]

    Crossover: 3d scene cross-modal alignment

    Sayan Deb Sarkar, Ondrej Miksik, Marc Pollefeys, Daniel Barath, and Iro Armeni. Crossover: 3d scene cross-modal alignment. InCVPR, 2025. 2

  60. [60]

    Learning structured representations with hyperbolic embed- dings.NeurIPS, 2024

    Aditya Sinha, Siqi Zeng, Makoto Yamada, and Han Zhao. Learning structured representations with hyperbolic embed- dings.NeurIPS, 2024. 3

  61. [61]

    Poincar\’e glove: Hyperbolic word embeddings

    Alexandru Tifrea, Gary B ´ecigneul, and Octavian-Eugen Ganea. Poincar\’e glove: Hyperbolic word embeddings. arXiv preprint arXiv:1810.06546, 2018. 3

  62. [62]

    Training data-efficient image transformers & distillation through at- tention

    Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herv ´e J´egou. Training data-efficient image transformers & distillation through at- tention. InICML, 2021. 10 24

  63. [63]

    A picture is worth more than 77 text tokens: Evaluating clip- style models on dense captions

    Jack Urbanek, Florian Bordes, Pietro Astolfi, Mary Williamson, Vasu Sharma, and Adriana Romero-Soriano. A picture is worth more than 77 text tokens: Evaluating clip- style models on dense captions. InCVPR, 2024. 7, 8, 11

  64. [64]

    Attention is all you need.NeurIPS, 2017

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need.NeurIPS, 2017. 10

  65. [65]

    Order-embeddings of images and language.arXiv preprint arXiv:1511.06361, 2015

    Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun. Order-embeddings of images and language.arXiv preprint arXiv:1511.06361, 2015. 1, 3

  66. [66]

    Sclip: Rethinking self-attention for dense vision-language inference

    Feng Wang, Jieru Mei, and Alan Yuille. Sclip: Rethinking self-attention for dense vision-language inference. InECCV,

  67. [67]

    Generalisation of structural knowl- edge in the hippocampal-entorhinal system.NeurIPS, 2018

    James Whittington, Timothy Muller, Shirely Mark, Caswell Barry, and Tim Behrens. Generalisation of structural knowl- edge in the hippocampal-entorhinal system.NeurIPS, 2018. 1

  68. [68]

    Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations

    Hao Wu, Jiayuan Mao, Yufeng Zhang, Yuning Jiang, Lei Li, Weiwei Sun, and Wei-Ying Ma. Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations. InCVPR, 2019. 2

  69. [69]

    Learning structure from the ground up—hierarchical representation learning by chunking.NeurIPS, 2022

    Shuchen Wu, No ´emi ´Elteto, Ishita Dasgupta, and Eric Schulz. Learning structure from the ground up—hierarchical representation learning by chunking.NeurIPS, 2022. 1

  70. [70]

    Fg- clip: Fine-grained visual and textual alignment.ICML, 2025

    Chunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li, Dawei Liang, Gengshen Zhang, Dawei Leng, and Yuhui Yin. Fg- clip: Fine-grained visual and textual alignment.ICML, 2025. 13

  71. [71]

    Empirical evaluation of rectified activations in convolutional network

    Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li. Empirical evaluation of rectified activations in convolutional network. arXiv preprint arXiv:1505.00853, 2015. 14

  72. [72]

    Hyp-uml: Hyper- bolic image retrieval with uncertainty-aware metric learning

    Shiyang Yan, Zongxuan Liu, and Lin Xu. Hyp-uml: Hyper- bolic image retrieval with uncertainty-aware metric learning. arXiv preprint arXiv:2310.08390, 2023. 2, 3, 4

  73. [73]

    Peter Young, Alice Lai, Micah Hodosh, and Julia Hocken- maier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.Transactions of the association for computational lin- guistics, 2014. 6, 10 25