Pith. sign in

REVIEW 2 major objections 2 minor 56 references

Gaia DR3 quasar proper motions set a detectable stochastic GW strain floor of order 10^{-11} below ~5.6 nHz.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-07-13 19:55 UTC pith:AQDX4JP6

load-bearing objection Abstract-only Gaia DR3 nHz GW forecast (~10^{-11} strain); full text is the wrong paper (GLA-CLIP), so the claim is unverifiable here. the 2 major comments →

arxiv 2603.23028 v3 pith:AQDX4JP6 submitted 2026-03-24 astro-ph.CO astro-ph.IMgr-qc

Low-Frequency Stochastic Gravitational-Wave Background in Gaia DR3 catalog

classification astro-ph.CO astro-ph.IMgr-qc
keywords gravitational wavesGaia DR3quasar proper motionsstochastic backgroundHellings-Downsvector spherical harmonicsnanohertzastrometry
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks whether the tiny, correlated deflections that a low-frequency stochastic gravitational-wave background would leave on the proper motions of distant quasars can be recovered from Gaia’s astrometric catalog. Using the real sky positions and published DR3 uncertainties of those quasars, the authors inject simulated GW signals and then try to extract them with two standard tools: a Vector Spherical Harmonic decomposition of the proper-motion field and the angular correlation function known as the Hellings–Downs curve. From the recovery statistics they derive a concrete sensitivity floor: with present DR3 errors a stochastic strain of roughly 10^{-11} is the lowest amplitude that could be detected when the spectrum is integrated over all frequencies below about 5.6 nHz (half the inverse of the 34-month DR3 time span). The same exercise forecasts that DR4, keeping the same quasars, could push the floor down to about 3×10^{-12}. The work therefore supplies a quantitative, data-driven bound on how far Gaia can already probe the nanohertz gravitational-wave sky.

Core claim

With the actual Gaia DR3 quasar sample, measurement errors, and sky coverage, the lowest stochastic gravitational-wave strain that can be detected is of order 10^{-11} for a spectrum integrated over frequencies ≲ 5.6 nHz; the same quasars observed through DR4 would improve that floor to roughly 3×10^{-12}.

What carries the argument

Simulated GW-induced proper-motion vectors are recovered by two complementary estimators—Vector Spherical Harmonic (VSH) decomposition of the vector field and the Hellings–Downs angular correlation function—applied to the real DR3 quasar positions and uncertainties; the detection threshold is read off from the recovery statistics of these estimators.

Load-bearing premise

The simulations that inject GW signals into the published DR3 proper-motion errors and sky distribution already capture every noise source that would limit a real search, so residual attitude, calibration, or quasar-structure systematics do not raise the true floor.

What would settle it

A real search performed on the actual Gaia DR3 proper-motion catalog that either recovers a Hellings–Downs or VSH signal at strain ≲ 10^{-11} or else returns a significantly higher upper limit would falsify the quoted sensitivity.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Existing Gaia DR3 data already place a quantitative upper bound on the integrated stochastic GW background below ~5.6 nHz.
  • The next data release (DR4) is forecast to improve that bound by a factor of roughly three without any increase in quasar numbers.
  • VSH analysis is more robust to uneven sky sampling and anisotropic noise than the pairwise Hellings–Downs estimator.
  • Different weighting and sky-restriction schemes can be traded against one another to optimize the final strain limit.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If residual systematics prove smaller than assumed, the same pipeline could already be competitive with early pulsar-timing-array bounds in the overlapping nanohertz band.
  • The method supplies an independent, purely astrometric cross-check that future multi-messenger GW catalogs can use to validate or refute a claimed stochastic background.
  • Extending the same simulation framework to Gaia’s full mission lifetime would map how the strain floor scales with baseline and with the eventual increase in quasar numbers.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The submitted abstract claims a forecast of the sensitivity of Gaia DR3 quasar proper motions to a low-frequency stochastic gravitational-wave background, using realistic sky positions and published uncertainties, recovered via Vector Spherical Harmonics (VSH) and Hellings–Downs angular correlations (HDC). It reports a detectable strain floor of order 10^{-11} for DR3 (improving to ~3×10^{-12} for DR4 with the same quasars) for a spectrum integrated below ~5.6 nHz (half the inverse of the 34-month DR3 timespan), and compares the robustness of VSH versus HDC under uneven sampling and anisotropic noise. The body of the manuscript, however, is an entirely different work (GLA-CLIP: a training-free open-vocabulary semantic segmentation method based on CLIP with key-value extension, proxy anchors, and dynamic normalization). No equations, simulation design, recovery statistics, noise model, or systematics treatment for the Gaia GW analysis appear in the supplied full text.

Significance. If the abstract’s forecast were supported by a complete, reproducible analysis, a Gaia-based nHz strain floor of ~10^{-11} would be a useful complementary constraint to pulsar-timing arrays in a sparsely covered band, and a careful VSH/HDC comparison under realistic sampling would be of methodological interest to the astrometric GW community. Those claims cannot be assessed from the manuscript as provided, because the scientific content of the body does not correspond to the title or abstract.

major comments (2)
  1. Title/abstract versus full text: the manuscript body is the GLA-CLIP computer-vision paper (training-free OVSS with sliding windows, Key-Value Extension, Proxy Anchor, Dynamic Normalization; Tables 1–12, Figs. 1–13, Eqs. (1)–(17) of that work). It contains no Gaia DR3 catalog analysis, no GW proper-motion imprint model, no VSH or Hellings–Downs estimators, and no strain-limit derivation. The central claims of the abstract (strain floor ~10^{-11} for DR3, ~3×10^{-12} for DR4, f ≲ 5.6 nHz) are therefore unsupported by any equation, figure, table, or error budget in the supplied text and cannot be refereed.
  2. Because the body is unrelated, the load-bearing premise of the abstract—that simulations using actual DR3 positions and published proper-motion uncertainties capture the dominant noise and systematics (attitude, calibration, quasar structure) that set a real search floor—cannot be examined. No residual systematics budget, no recovery-to-strain mapping, and no comparison of VSH vs HDC under anisotropic noise are present. This is not a presentation issue; the scientific content required to support the abstract is missing.
minor comments (2)
  1. Even if the correct Gaia GW manuscript were substituted, the abstract alone leaves free parameters (DR4 error scaling at fixed quasar count; data-restriction and weighting schemes) without quantitative definition; those would need explicit equations and tables in any resubmission.
  2. The supplied GLA-CLIP body has its own internal presentation issues (e.g., typographical noise in figure captions, ‘hyperaprater’ in §4.1, inconsistent crop/stride settings across tables) but they are irrelevant to the claimed Gaia GW science and do not affect the recommendation.

Circularity Check

0 steps flagged

No circularity found: Gaia GW strain floor is a simulation-based sensitivity forecast; supplied body is an unrelated CV paper so no load-bearing reduction can be exhibited.

full rationale

The claimed result (detectable stochastic GW strain ~10^{-11} for Gaia DR3 proper motions, ~3e-12 for DR4, integrated below ~5.6 nHz) is stated only in the abstract of arXiv:2603.23028 as a forecast from forward simulations that inject GW imprints into real quasar sky positions and published measurement uncertainties, then recover with VSH and Hellings–Downs estimators. That construction is a standard sensitivity study: the strain amplitude is an input to the simulation, and the reported floor is the recovery threshold under those noise assumptions—not a quantity defined to equal a fitted parameter by construction, nor a uniqueness claim imported from self-citation. The CACHEABLE full manuscript body is an unrelated computer-vision paper (GLA-CLIP), so no equations, recovery statistics, or systematics model from the Gaia analysis are available to quote. Per the hard rules, circularity may be claimed only when a specific reduction can be exhibited from the paper’s own text; none can. Therefore score 0 with empty steps. (Any concern that residual attitude/calibration systematics are under-modeled is a correctness/completeness risk, not circularity.)

Axiom & Free-Parameter Ledger

3 free parameters · 3 axioms · 0 invented entities

Abstract-only review of a simulation forecast. Load-bearing inputs are the Gaia DR3 timespan, quasar sky distribution, published proper-motion uncertainties, the assumption that GW effects appear as correlated proper motions recoverable by VSH/HDC, and the frequency cutoff at half the inverse mission length. No free parameters are numerically fitted in the abstract; the strain numbers are simulation outputs. Invented entities are none—the GW background and estimators are standard.

free parameters (3)
  • DR3 observational timespan (34 months) and implied frequency cutoff (~5.6 nHz) = 34 months; ~5.6 nHz
    Sets the integrated frequency band for the quoted strain limit; taken as mission property but directly defines the claim's domain.
  • Assumed DR4 proper-motion error improvement (same quasar count) = strain floor ~3e-12
    Used to quote the improved ~3×10^{-12} floor; abstract does not detail the error model.
  • Data-restriction and weighting schemes
    Abstract states these influence the final strain estimate; specific choices are free analysis knobs.
axioms (3)
  • domain assumption A stochastic GW background imprints a recoverable correlated proper-motion pattern on distant quasars that can be extracted by VSH and/or Hellings-Downs estimators.
    Core physical and statistical premise of the entire forecast (abstract).
  • domain assumption Published Gaia DR3 quasar positions and proper-motion uncertainties, plus the actual sky distribution, are adequate inputs for a realistic sensitivity simulation.
    Simulation design premise; systematics beyond catalog errors are only mentioned as evaluated, not specified.
  • domain assumption The relevant GW spectrum can be treated as integrated over all frequencies below half the inverse observational timespan.
    Defines the ~5.6 nHz band of the quoted limit.

reviewed 2026-07-13 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Low-Frequency Stochastic Gravitational-Wave Background in Gaia DR3 catalog." pith.science (2026). https://pith.science/paper/AQDX4JP6

@misc{pith2026260323028,
  author       = {Pith},
  title        = {Pith review of: Low-Frequency Stochastic Gravitational-Wave Background in Gaia DR3 catalog},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AQDX4JP6}},
  note         = {Machine review of arXiv:2603.23028}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We investigate the potential to detect low-frequency gravitational waves (GWs) through their imprints on the proper motions of distant quasars observed by the Gaia mission. Using astrometric data from Gaia DR3, we simulate the effect of GWs on the proper motions of quasars, incorporating their actual sky positions and measurement uncertainties. We investigate two closely related data analysis techniques for the extraction and characterization of GW signals from quasar proper motions: Vector Spherical Harmonics (VSH) and angular correlation functions, commonly referred to as Hellings-Downs curves (HDC). Using realistic simulated data, we forecast their sensitivity and accuracy to GWs, and evaluate the impact of systematic errors. From these simulations, we derive an upper limit on the amplitude of a stochastic GW background, constrained by the observational timespan, astrometric precision, and the sky distribution of quasars. VSH decomposition appears less sensitive to uneven sky sampling and anisotropic noise. The HDC approach retains a larger fraction of the pairwise correlation information and therefore exhibits higher raw statistical sensitivity under idealized conditions. We find that, with Gaia DR3 proper motion errors, the lower limit for a detectable GW strain is of 10^{-11}, with possible improvements to about 3 x 10^{-12} for the next Gaia Data Release 4 (for the same number of quasars). This limit holds for a stochastic GW spectrum integrated over all frequencies less than half the inverse of the 34-month observational timespan of Gaia DR3, corresponding to approximately 5.6 nHz. We also investigate how different data-restriction and weighting schemes influence the final estimate of the gravitational wave strain.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

56 extracted references · 5 linked inside Pith

  1. [1]

    Self-calibrated clip for training-free open-vocabulary segmentation.arXiv preprint arXiv:2411.15869, 2024

    Sule Bai, Yong Liu, Yifei Han, Haoji Zhang, and Yansong Tang. Self-calibrated clip for training-free open-vocabulary segmentation.arXiv preprint arXiv:2411.15869, 2024. 3

  2. [2]

    Floss: Free lunch in open-vocabulary semantic segmentation.Proceedings of the IEEE/CVF International Conference on Computer Vision,

    Yasser Benigmim, Mohammad Fahes, Tuan-Hung Vu, An- drei Bursuc, and Raoul de Charette. Floss: Free lunch in open-vocabulary semantic segmentation.Proceedings of the IEEE/CVF International Conference on Computer Vision,

  3. [3]

    Grounding everything: Emerging localiza- tion properties in vision-language transformers

    Walid Bousselham, Felix Petersen, Vittorio Ferrari, and Hilde Kuehne. Grounding everything: Emerging localiza- tion properties in vision-language transformers. InProceed- ings of the IEEE conference on computer vision and pattern recognition, pages 3828–3837, 2024. 1, 2

  4. [4]

    Coco- stuff: Thing and stuff classes in context

    Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Coco- stuff: Thing and stuff classes in context. InProceedings of the IEEE conference on computer vision and pattern recog- nition, pages 1209–1218, 2018. 5, 6, 7, 1, 2

  5. [5]

    Vdd: Varied drone dataset for semantic segmentation.Journal of Visual Communication and Image Representation, 109:104429, 2025

    Wenxiao Cai, Ke Jin, Jinyan Hou, Cong Guo, Letian Wu, and Wankou Yang. Vdd: Varied drone dataset for semantic segmentation.Journal of Visual Communication and Image Representation, 109:104429, 2025. 4

  6. [6]

    Emerg- ing properties in self-supervised vision transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. InPro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9650–9660, 2021. 3

  7. [7]

    Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs

    Junbum Cha, Jonghwan Mun, and Byungseok Roh. Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs. InProceedings of the IEEE conference on computer vision and pattern recog- nition, pages 11165–11174, 2023. 6

  8. [8]

    Training-free class purification for open-vocabulary semantic segmentation

    Qi Chen, Lingxiao Yang, Yun Chen, Nailong Zhao, Jian- huang Lai, Jie Shao, and Xiaohua Xie. Training-free class purification for open-vocabulary semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 23124–23134, 2025. 5, 6

  9. [9]

    Masked-attention mask transformer for universal image segmentation

    Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexan- der Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. InProceed- ings of the IEEE conference on computer vision and pattern recognition, pages 1290–1299, 2022. 2

  10. [10]

    Plug-in feedback self-adaptive attention in clip for training- free open-vocabulary segmentation

    Zhixiang Chi, Yanan Wu, Li Gu, Huan Liu, Ziqiang Wang, Yang Zhang, Yang Wang, and Konstantinos Plataniotis. Plug-in feedback self-adaptive attention in clip for training- free open-vocabulary segmentation. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 22815–22825, 2025. 6

  11. [11]

    Cat-seg: Cost aggregation for open-vocabulary semantic segmenta- tion

    Seokju Cho, Heeseong Shin, Sunghwan Hong, Anurag Arnab, Paul Hongsuck Seo, and Seungryong Kim. Cat-seg: Cost aggregation for open-vocabulary semantic segmenta- tion. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 4113–4123, 2024. 2

  12. [12]

    The cityscapes dataset for semantic urban scene understanding

    Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. InProceed- ings of the IEEE conference on computer vision and pattern recognition, pages 3213–3223, 2016. 5, 6, 7, 1, 3, 4

  13. [13]

    Vision transformers need registers.arXiv preprint arXiv:2309.16588, 2023

    Timoth ´ee Darcet, Maxime Oquab, Julien Mairal, and Pi- otr Bojanowski. Vision transformers need registers.arXiv preprint arXiv:2309.16588, 2023. 8, 2

  14. [14]

    De- coupling zero-shot semantic segmentation

    Jian Ding, Nan Xue, Gui-Song Xia, and Dengxin Dai. De- coupling zero-shot semantic segmentation. InProceedings of the IEEE conference on computer vision and pattern recog- nition, pages 11583–11592, 2022. 2

  15. [15]

    Open- vocabulary universal image segmentation with maskclip

    Zheng Ding, Jieke Wang, and Zhuowen Tu. Open- vocabulary universal image segmentation with maskclip. arXiv preprint arXiv:2208.08984, 2022. 2

  16. [16]

    Dih-clip: Unleashing the diversity of multi-head self-attention for training-free open-vocabulary semantic segmentation

    Songsong Duan, Xi Yang, and Nannan Wang. Dih-clip: Unleashing the diversity of multi-head self-attention for training-free open-vocabulary semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 22794–22803, 2025. 6

  17. [17]

    The pascal visual object classes challenge 2012 (voc2012) results (2012), 2011

    Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes challenge 2012 (voc2012) results (2012), 2011. 5, 6, 7, 1, 2, 4

  18. [18]

    Pay atten- tion to your neighbours: Training-free open-vocabulary se- mantic segmentation

    Sina Hajimiri, Ismail Ben Ayed, and Jose Dolz. Pay atten- tion to your neighbours: Training-free open-vocabulary se- mantic segmentation. InProceedings of the IEEE/CVF Win- ter Conference on Applications of Computer Vision, pages 5061–5071. IEEE, 2025. 1, 2, 3, 6

  19. [19]

    Scaling up visual and vision-language representa- tion learning with noisy text supervision

    Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representa- tion learning with noisy text supervision. InInternational conference on machine learning, pages 4904–4916. PMLR,

  20. [20]

    In defense of lazy visual grounding for open-vocabulary semantic segmentation

    Dahyun Kang and Minsu Cho. In defense of lazy visual grounding for open-vocabulary semantic segmentation. In European Conference on Computer Vision, pages 143–164. Springer, 2024. 6

  21. [21]

    Diffusion models for zero-shot open-vocabulary segmentation.arXiv e-prints, pages arXiv–2306, 2023

    Laurynas Karazija, Iro Laina, Andrea Vedaldi, and Christian Rupprecht. Diffusion models for zero-shot open-vocabulary segmentation.arXiv e-prints, pages arXiv–2306, 2023. 2

  22. [22]

    Distilling spectral graph for object-context aware open-vocabulary semantic segmenta- tion

    Chanyoung Kim, Dayun Ju, Woojung Han, Ming-Hsuan Yang, and Seong Jae Hwang. Distilling spectral graph for object-context aware open-vocabulary semantic segmenta- tion. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 15033–15042, 2025. 2, 5, 6, 7, 1, 4

  23. [23]

    Clearclip: Decom- posing clip representations for dense vision-language infer- ence

    Mengcheng Lan, Chaofeng Chen, Yiping Ke, Xinjiang Wang, Litong Feng, and Wayne Zhang. Clearclip: Decom- posing clip representations for dense vision-language infer- ence. InEuropean Conference on Computer Vision, pages 143–160. Springer, 2024. 1, 3, 6, 2

  24. [24]

    Proxyclip: Proxy attention improves clip for open-vocabulary segmentation

    Mengcheng Lan, Chaofeng Chen, Yiping Ke, Xinjiang Wang, Litong Feng, and Wayne Zhang. Proxyclip: Proxy attention improves clip for open-vocabulary segmentation. InEuropean Conference on Computer Vision, pages 70–88. Springer, 2024. 1, 2, 3, 5, 6, 7, 4

  25. [25]

    Segearth-ov: Towards training-free open-vocabulary segmentation for remote sens- ing images

    Kaiyu Li, Ruixun Liu, Xiangyong Cao, Xueru Bai, Feng Zhou, Deyu Meng, and Zhi Wang. Segearth-ov: Towards training-free open-vocabulary segmentation for remote sens- ing images. InProceedings of the IEEE conference on com- puter vision and pattern recognition, pages 10545–10556,

  26. [26]

    Clip surgery for better explainability with enhancement in open- vocabulary tasks.arXiv e-prints, pages arXiv–2304, 2023

    Yi Li, Hualiang Wang, Yiqun Duan, and Xiaomeng Li. Clip surgery for better explainability with enhancement in open- vocabulary tasks.arXiv e-prints, pages arXiv–2304, 2023. 2

  27. [27]

    Open-vocabulary semantic segmentation with mask-adapted clip

    Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu. Open-vocabulary semantic segmentation with mask-adapted clip. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 7061– 7070, 2023. 2

  28. [28]

    Unveiling the knowledge of clip for training-free open-vocabulary semantic segmentation

    Yajie Liu, Guodong Wang, Jinjin Zhang, Qingjie Liu, and Di Huang. Unveiling the knowledge of clip for training-free open-vocabulary semantic segmentation. InProceedings of the AAAI Conference on Artificial Intelligence, pages 5649– 5657, 2025. 1

  29. [29]

    Image segmenta- tion using text and image prompts

    Timo L ¨uddecke and Alexander Ecker. Image segmenta- tion using text and image prompts. InProceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 7086–7096, 2022. 2

  30. [30]

    Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmentation

    Huaishao Luo, Junwei Bao, Youzheng Wu, Xiaodong He, and Tianrui Li. Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmentation. InIn- ternational conference on machine learning, pages 23033– 23044. PMLR, 2023. 2

  31. [31]

    Uavid: A semantic segmentation dataset for uav imagery.ISPRS journal of photogrammetry and remote sensing, 165:108–119, 2020

    Ye Lyu, George V osselman, Gui-Song Xia, Alper Yilmaz, and Michael Ying Yang. Uavid: A semantic segmentation dataset for uav imagery.ISPRS journal of photogrammetry and remote sensing, 165:108–119, 2020. 4

  32. [32]

    The role of context for object detection and semantic segmentation in the wild

    Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, and Alan Yuille. The role of context for object detection and semantic segmentation in the wild. InProceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 891–898, 2014. 5, 1, 2, 3, 4, 6

  33. [33]

    Dinov2: Learning robust visual features without supervision

    Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 8

  34. [34]

    Learning transferable visual models from natural language supervi- sion

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. InInternational conference on machine learning, pages 8748–8763. PmLR, 2021. 1, 2, 3

  35. [35]

    Denseclip: Language-guided dense prediction with context- aware prompting

    Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu. Denseclip: Language-guided dense prediction with context- aware prompting. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 18082– 18091, 2022. 2

  36. [36]

    Progressive proxy anchor propagation for unsupervised semantic segmentation

    Hyun Seok Seong, WonJun Moon, SuBeen Lee, and Jae-Pil Heo. Progressive proxy anchor propagation for unsupervised semantic segmentation. InEuropean Conference on Com- puter Vision, pages 472–490. Springer, 2024. 4

  37. [37]

    Open-vocabulary semantic segmentation with image embedding balancing

    Xiangheng Shan, Dongyue Wu, Guilin Zhu, Yuanjie Shao, Nong Sang, and Changxin Gao. Open-vocabulary semantic segmentation with image embedding balancing. InProceed- ings of the IEEE conference on computer vision and pattern recognition, pages 28412–28421, 2024. 2

  38. [38]

    Ex- plore the potential of clip for training-free open vocabulary semantic segmentation

    Tong Shao, Zhuotao Tian, Hang Zhao, and Jingyong Su. Ex- plore the potential of clip for training-free open vocabulary semantic segmentation. InEuropean Conference on Com- puter Vision, pages 139–156. Springer, 2024. 1, 2

  39. [39]

    Har- nessing vision foundation models for high-performance, training-free open vocabulary segmentation.arXiv preprint arXiv:2411.09219, 2024

    Yuheng Shi, Minjing Dong, and Chang Xu. Har- nessing vision foundation models for high-performance, training-free open vocabulary segmentation.arXiv preprint arXiv:2411.09219, 2024. 3

  40. [40]

    Reco: Retrieve and co-segment for zero-shot transfer.Advances in neural information processing systems, 35:33754–33767,

    Gyungin Shin, Weidi Xie, and Samuel Albanie. Reco: Retrieve and co-segment for zero-shot transfer.Advances in neural information processing systems, 35:33754–33767,

  41. [41]

    Dinov3.arXiv preprint arXiv:2508.10104, 2025

    Oriane Sim ´eoni, Huy V V o, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Micha ¨el Ramamonjisoa, et al. Dinov3.arXiv preprint arXiv:2508.10104, 2025. 8

  42. [42]

    Sclip: Rethink- ing self-attention for dense vision-language inference

    Feng Wang, Jieru Mei, and Alan Yuille. Sclip: Rethink- ing self-attention for dense vision-language inference. In European Conference on Computer Vision, pages 315–332. Springer, 2024. 1, 2, 6, 4

  43. [43]

    Sclip: Rethink- ing self-attention for dense vision-language inference

    Feng Wang, Jieru Mei, and Alan Yuille. Sclip: Rethink- ing self-attention for dense vision-language inference. In European Conference on Computer Vision, pages 315–332. Springer, 2024. 2

  44. [44]

    Clip-diy: Clip dense infer- ence yields open-vocabulary semantic segmentation for-free

    Monika Wysocza ´nska, Micha ¨el Ramamonjisoa, Tomasz Trzci´nski, and Oriane Sim ´eoni. Clip-diy: Clip dense infer- ence yields open-vocabulary semantic segmentation for-free. InProceedings of the IEEE/CVF Winter Conference on Ap- plications of Computer Vision, pages 1403–1413, 2024. 6

  45. [45]

    Clip-dinoiser: Teaching clip a few dino tricks for open- vocabulary semantic segmentation

    Monika Wysocza ´nska, Oriane Sim´eoni, Micha¨el Ramamon- jisoa, Andrei Bursuc, Tomasz Trzci ´nski, and Patrick P ´erez. Clip-dinoiser: Teaching clip a few dino tricks for open- vocabulary semantic segmentation. InEuropean Conference on Computer Vision, pages 320–337. Springer, 2024. 6, 1, 4

  46. [46]

    Groupvit: Semantic segmentation emerges from text supervision

    Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang. Groupvit: Semantic segmentation emerges from text supervision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 18134–18144, 2022. 6, 1

  47. [47]

    Open-vocabulary panop- tic segmentation with text-to-image diffusion models

    Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon, Xiao- long Wang, and Shalini De Mello. Open-vocabulary panop- tic segmentation with text-to-image diffusion models. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2955–2966, 2023. 2

  48. [48]

    A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- language model

    Mengde Xu, Zheng Zhang, Fangyun Wei, Yutong Lin, Yue Cao, Han Hu, and Xiang Bai. A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- language model. InEuropean Conference on Computer Vi- sion, pages 736–753. Springer, 2022. 2

  49. [49]

    San: side adapter network for open-vocabulary se- mantic segmentation.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(12):15546–15561, 2023

    Mengde Xu, Zheng Zhang, Fangyun Wei, Han Hu, and Xi- ang Bai. San: side adapter network for open-vocabulary se- mantic segmentation.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(12):15546–15561, 2023. 2

  50. [50]

    Resclip: Residual attention for training-free dense vision- language inference

    Yuhang Yang, Jinhong Deng, Wen Li, and Lixin Duan. Resclip: Residual attention for training-free dense vision- language inference. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 29968–29978,

  51. [51]

    Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip.Advances in neural information processing systems, 36:32215–32234,

    Qihang Yu, Ju He, Xueqing Deng, Xiaohui Shen, and Liang- Chieh Chen. Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip.Advances in neural information processing systems, 36:32215–32234,

  52. [52]

    Semantic under- standing of scenes through the ade20k dataset.International Journal of Computer Vision, 127(3):302–321, 2019

    Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fi- dler, Adela Barriuso, and Antonio Torralba. Semantic under- standing of scenes through the ade20k dataset.International Journal of Computer Vision, 127(3):302–321, 2019. 5, 1, 2, 4

  53. [53]

    Extract free dense labels from clip

    Chong Zhou, Chen Change Loy, and Bo Dai. Extract free dense labels from clip. InEuropean Conference on Com- puter Vision, pages 696–712. Springer, 2022. 1, 2

  54. [54]

    Zegclip: Towards adapting clip for zero-shot se- mantic segmentation

    Ziqin Zhou, Yinjie Lei, Bowen Zhang, Lingqiao Liu, and Yifan Liu. Zegclip: Towards adapting clip for zero-shot se- mantic segmentation. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 11175– 11185, 2023. 2 Looking Beyond the Window: Global-Local Aligned CLIP for Training-free Open-Vocabulary Semantic Segmentation Sup...

  55. [55]

    background

    datasets. We adopt a sliding-window strategy with a 336×336 window and 112×112 stride. For the back- ground class, rather than directly using the text prompt “background”, we employ a renaming strategy in which multiple class names associated with background semantics are grouped and used as substitutes. This follows the official ProxyCLIP implementation....

  56. [56]

    is assigned to 0.2, 0.15, 0.25, for each. B.4. SCLIP Setting In SCLIP [42] setting in Tab 1, We resize input images with a short side of 336 and perform sliding-window inference with a 224×224 window and 112 stride. Only for the Cityscapes, we resize the short side of 560. We use rename trick on both background class and several other classes, following o...

This paper was first reviewed by grok-4.5 on July 13, 2026.