Pith. sign in

REVIEW 2 major objections 2 minor 2 cited by

Peel neighborhoods

T0 review · 2 major / 2 minor · reviewed 2026-07-14 · grok-4.5

Pith's one-line read Peel neighborhoods give a canonical, parameter-free local neighborhood in finite metric spaces of strict negative type.

desk verdict Only the abstract of Peel neighborhoods is available; the supplied full text is a mismatched CV paper, so the math cannot be audited—but the claimed construction is a clean, classical-hypothesis idea that would be useful if the body delivers. read the letter →

arxiv 2603.26645 v2 pith:NI6SDO5A submitted 2026-03-27 math.MG math.AT

classification math.MGmath.AT MSC 51F9954E3562H30
keywords peelneighborhoodsstrictnegativetypefinitemetricspaceslocaldimensionsingularitiesstratifiedmanifoldsgeometry
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper defines peel neighborhoods as a canonical way to pick local neighborhoods in a finite metric space that has strict negative type. The construction needs no free parameters, can be computed efficiently, and scales when a soft upper bound is put on radius or size. That scaling makes peel neighborhoods practical for describing geometry and topology at a fine scale. As a concrete use, they support fast local-dimension estimates and singularity detection on point samples from stratified manifolds. A sympathetic reader cares because many data sets are finite metric spaces, and a parameter-free local neighborhood is a basic building block for geometric and topological analysis of those spaces.

What carries the argument

Peel neighborhoods: a canonical, parameter-free local neighborhood construction in finite metric spaces of strict negative type, made scalable by a soft threshold on radius or cardinality.

What would settle it

Take a finite metric space known not to be of strict negative type, or a stratified-manifold sample whose metric fails the condition, and check whether peel neighborhoods remain well-defined, unique, efficiently computable, and still recover local dimension and singularities as claimed.

Watch

Extended reading notes

Core claim

In a finite metric space of strict negative type there exists a canonical, parameter-free, efficiently computable notion of local neighborhood—the peel neighborhood—and, once a soft threshold bounds its radius or cardinality, these neighborhoods can be computed at scale and used for microscopic geometric and topological descriptions, including local-dimension estimation and singularity detection on stratified-manifold samples.

Load-bearing premise

The finite metric space must be of strict negative type; if real data metrics fail that classical condition, the claimed canonicity and algorithms may not apply as stated.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The abstract claims to introduce peel neighborhoods: a canonical, parameter-free, efficiently computable local neighborhood construction on finite metric spaces of strict negative type, optionally capped by a soft radius/cardinality threshold for scalability, with applications to local-dimension estimation and singularity detection on samples from stratified manifolds. The supplied full manuscript body, however, is an entirely different paper (EgoPoint-Ground / SV-CoT on egocentric hand-pointing visual grounding, arXiv:2603.26646, cs.CV). No definitions, theorems, algorithms, complexity statements, or experiments for peel neighborhoods appear in the body.

Significance. If the abstract’s claims hold for a correctly supplied math.MG manuscript, a canonical parameter-free neighborhood notion with scalable computation and concrete geometric applications would be of genuine interest in metric geometry and geometric data analysis. As submitted, that significance cannot be assessed: the body contains no mathematical content matching the title or abstract, so no credit can be given for proofs, complexity bounds, or empirical performance.

major comments (2)
  1. Title/abstract vs. body mismatch: paper_id 2603.26645 and the abstract concern peel neighborhoods in finite metric spaces of strict negative type, but the full manuscript text is the unrelated EgoPoint-Ground CV paper (hand-pointing referring expressions, SV-CoT, Tables 1–4, etc.). No definition of peel neighborhood, no negative-type hypothesis, no algorithm, and no local-dimension/singularity results are present. The central claims are therefore unverifiable from the submission as provided.
  2. Because the mathematical body is absent, load-bearing claims in the abstract—canonicity, parameter-freeness (beyond the acknowledged soft threshold), efficient computability, and utility for local dimension and singularity detection—cannot be checked against definitions, proofs, or experiments. A referee report on the actual math.MG contribution is not possible until the correct manuscript is supplied.
minor comments (2)
  1. The abstract alone is coherent on its face and conditions the construction on the classical hypothesis of strict negative type; residual circularity risk from the soft threshold is low as stated. These points cannot be elevated to major technical objections without the correct body.
  2. If the wrong full text was attached in error, the authors should resubmit the actual peel-neighborhoods manuscript; the present package is not reviewable as a math.MG paper.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity detectable: only the abstract of Peel neighborhoods is available; the supplied full text is a mismatched CV paper.

full rationale

The target paper (arXiv:2603.26645, Peel neighborhoods) is represented only by its abstract. That abstract defines peel neighborhoods as a canonical, parameter-free construction on finite metric spaces of strict negative type, with an optional soft threshold used solely as a computational cap on radius or cardinality, not as a fitted parameter that defines the neighborhoods or the quantities they estimate. Local-dimension estimates and singularity detection are presented as downstream applications, not as inputs to the definition. No equations, self-citations, uniqueness theorems, or fitted-then-predicted quantities appear in the available text, so no reduction of a claimed prediction to its own inputs can be exhibited. The CACHEABLE full manuscript is a completely different paper (EgoPoint-Ground / SV-CoT, arXiv:2603.26646); it cannot be used to audit the derivation chain of 2603.26645. Under the abstract alone the construction is self-contained and non-circular; residual risk that the missing mathematical body hides definitional circularity cannot be audited and is not scored as circularity.

Assumptions & free parameters 1 free parameters · 2 assumptions · 1 invented entities

Abstract-only review of peel neighborhoods. Load-bearing background is the classical metric condition of strict negative type and the existence of a well-defined finite metric on the sample. Soft thresholds are optional computational parameters, not definitional free parameters of the neighborhood itself. No invented physical entities appear. Full axiom inventory would require the missing mathematical body.

free parameters (1)
  • soft threshold on radius or cardinality
    Abstract states a soft threshold is used to upper-bound radius or cardinality for scalability; if this threshold is chosen by hand or tuned to data, it is a free computational parameter even if the ideal peel neighborhood is parameter-free.
assumptions (2)
  • domain assumption The finite metric space is of strict negative type.
    The entire construction is stated only for spaces of strict negative type; this is a classical metric hypothesis, not proved in the abstract.
  • domain assumption Samples are drawn from stratified manifolds (for the application).
    Local-dimension and singularity claims are illustrated on stratified-manifold samples; the abstract assumes that modeling regime.
invented entities (1)
  • peel neighborhood
    purpose: Canonical local neighborhood for microscopic geometric/topological description in strict-negative-type finite metric spaces.
    Primary new object of the paper; independent evidence would be theorems, algorithms, and experiments in the body, which are not available here.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Peel neighborhoods." pith.science (2026). https://pith.science/paper/NI6SDO5A

@misc{pith2026260326645,
  author       = {Pith},
  title        = {Pith review of: Peel neighborhoods},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NI6SDO5A}},
  note         = {Machine review of arXiv:2603.26645}
}
read the original abstract

We introduce the canonical, parameter-free, and efficiently computable notion of peel neighborhoods in a finite metric space of strict negative type. Using a soft threshold to upper bound their radius or cardinality allows peel neighborhoods to be computed at scale, enabling useful microscopic descriptions of geometry and topology. As an example of their utility, peel neighborhoods enable efficient and performant estimates of local dimension and detections of singularities in samples from stratified manifolds.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Magnitude homology and Euler characteristics of directed acyclic graphs

    math.AT 2026-07 conditional novelty 5.0 of 10

    A decategorification shortcut turns magnitude-homology Euler characteristic computation for DAGs into a polynomial-arithmetic linear solve, with a proof-of-concept on MLP activation graphs.

  2. Scalably computing metric magnitude

    math.NA 2026-07 conditional novelty 5.0 of 10

    Hierarchical low-rank solvers beat dense and sparsified approaches for metric magnitude solves in experiments up to n=30,000, with a projected path to n≈10^5 via a containerized STRUMPACK/MPI pipeline.

Reference graph

Works this paper leans on

119 extracted references · 25 linked inside Pith · cited by 2 Pith papers

  1. [1]

    ArXiv Preprint ArXiv:2303.08774 (2023)

    Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al.: GPT-4 Technical Report. ArXiv Preprint ArXiv:2303.08774 (2023)

  2. [2]

    Ahuja, C.: Communication Beyond Words: Grounding Visual Body Motion with Language. Ph.D. thesis, Carnegie Mellon University (2022)

  3. [3]

    ArXiv Preprint ArXiv:2509.23661 (2025)

    An, X., Xie, Y., Yang, K., Zhang, W., Zhao, X., Cheng, Z., Wang, Y., Xu, S., Chen, C., Wu, C., et al.: Llava-onevision-1.5: Fully Open Framework for Democratized Multimodal Training. ArXiv Preprint ArXiv:2509.23661 (2025)

  4. [4]

    In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition

    Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I., Gould, S., Van Den Hengel, A.: Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 3674– 3683 (2018)

  5. [5]

    In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition

    Andriluka, M., Pishchulin, L., Gehler, P., Schiele, B.: 2D Human Pose Estimation: New Benchmark and State of the Art Analysis. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 3686–3693 (2014)

  6. [6]

    In: Proceedings of the IEEE International Conference on Computer Vision

    Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., Parikh, D.: VQA: Visual Question Answering. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 2425–2433 (2015)

  7. [7]

    Computers11, 28 (2022)

    Arena, F., Collotta, M., Pau, G., Termine, F.: An Overview of Augmented Reality. Computers11, 28 (2022)

  8. [8]

    ArXiv Preprint ArXiv:2309.16609 (2023)

    Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et al.: Qwen Technical Report. ArXiv Preprint ArXiv:2309.16609 (2023)

Show all 119 references
  1. [9]

    In: Proceedings of the European Conference on Computer Vision

    Bansal, A., Sikka, K., Sharma, G., Chellappa, R., Divakaran, A.: Zero-Shot Object Detection. In: Proceedings of the European Conference on Computer Vision. pp. 384–400 (2018)

  2. [10]

    In: Proceedings of the European Conference on Computer Vision

    Bansal, S., Arora, C., Jawahar, C.: My View is the Best View: Procedure Learning from Egocentric Videos. In: Proceedings of the European Conference on Computer Vision. pp. 657–675 (2022)

  3. [11]

    In: Healthcare

    Banyai, A.D., Bris�an, C.: Robotics in Physical Rehabilitation: Systematic Review. In: Healthcare. vol. 12, p. 1720 (2024)

  4. [12]

    In: European Conference on Computer Vision

    Bhattacharyya, P., Mitton, J., Page, R., Morgan, O., Menzies, B., Homewood, G., Jacobs, K., Baesso, P., Trickett, D., Mair, C., et al.: Helios: An extremely low power event-based gesture recognition for always-on smart eyewear. In: European Conference on Computer Vision. pp. 1...

  5. [13]

    Handbook of aug- mented reality pp

    Carmigniani, J., Furht, B.: Augmented Reality: An Overview. Handbook of aug- mented reality pp. 3–46 (2011)

  6. [14]

    International Journal of Computer Applications120(2015)

    Chandel, H., Vatta, S.: Occlusion Detection and Handling: A Review. International Journal of Computer Applications120(2015)

  7. [15]

    ArXiv Preprint ArXiv:2410.21091 (2024)

    Chen, J., Grubert, J., Kristensson, P.O.: Large Language Model-assisted Speech and Pointing Benefits Multiple 3D Object Selection in Virtual Reality. ArXiv Preprint ArXiv:2410.21091 (2024)

  8. [16]

    ArXiv Preprint ArXiv:2306.15195 (2023)

    Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., Zhao, R.: Shikra: Unleashing Multimodal LLM’s Referential Dialogue Magic. ArXiv Preprint ArXiv:2306.15195 (2023)

  9. [17]

    In: European Conference on Computer Vision

    Chen, W., Chen, L., Wu, Y.: An Efficient and Effective Transformer Decoder- Based Framework for Multi-Task Visual Grounding. In: European Conference on Computer Vision. pp. 125–141. Springer (2024) Abbreviated paper title 17

  10. [18]

    In: Proceedings of the IEEE/CVF International Conference on Computer Vision

    Chen, Y., Li, Q., Kong, D., Kei, Y.L., Zhu, S.C., Gao, T., Zhu, Y., Huang, S.: YouRefIt: Embodied Reference Understanding With Language and Gesture. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 1385–1395 (2021)

  11. [19]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Chen, Z., Wang, P., Ma, L., Wong, K.Y.K., Wu, Q.: Cops-Ref: A New Dataset and Task on Compositional Referring Expression Comprehension. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 10086–10095 (2020)

  12. [20]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Cheng, T., Song, L., Ge, Y., Liu, W., Wang, X., Shan, Y.: YOLO-World: Real-Time Open-Vocabulary Object Detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 16901–16911 (2024)

  13. [21]

    IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing17, 19640–19667 (2024)

    Chu, Y., Ye, M., Qian, Y.: Fine-Grained Image Recognition Methods and Their Applications in Remote Sensing Images: A Review. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing17, 19640–19667 (2024)

  14. [22]

    ArXiv Preprint ArXiv:2507.22062 (2025)

    Chuang, Y.S., Li, Y., Wang, D., Yeh, C.F., Lyu, K., Raghavendra, R., Glass, J., Huang, L., Weston, J., Zettlemoyer, L., et al.: Meta CLIP 2: A Worldwide Scaling Recipe. ArXiv Preprint ArXiv:2507.22062 (2025)

  15. [23]

    In: European conference on computer vision

    Constantin, S., Eyiokur, F.I., Yaman, D., Bärmann, L., Waibel, A.: Interactive Multimodal Robot Dialog Using Pointing Gesture Recognition. In: European conference on computer vision. pp. 640–657. Springer (2022)

  16. [24]

    In: Proceedings of the Advances in Neural Information Processing Systems

    Dai, M., Yang, L., Xu, Y., Feng, Z., Yang, W.: SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion. In: Proceedings of the Advances in Neural Information Processing Systems. pp. 121670–121698 (2024)

  17. [25]

    In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition

    De Vries, H., Strub, F., Chandar, S., Pietquin, O., Larochelle, H., Courville, A.: GuessWhat?! Visual Object Discovery Through Multi-Modal Dialogue. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 5503–5512 (2017)

  18. [26]

    arXiv preprint arXiv:2410.00253 (2024)

    Deichler, A., O’Regan, J., Beskow, J.: MM-Conv: A Multi-modal Conversational Dataset for Virtual Humans. arXiv preprint arXiv:2410.00253 (2024)

  19. [27]

    In: European Conference on Computer Vision

    Delmas, G., Weinzaepfel, P., Moreno-Noguer, F., Rogez, G.: PoseEmbroider: Towards a 3D, Visual, Semantic-aware Human Pose Representation. In: European Conference on Computer Vision. pp. 55–73. Springer (2024)

  20. [28]

    IEEE Transactions on Emerging Topics in Compu- tational Intelligence6, 230–244 (2022)

    Duan, J., Yu, S., Tan, H.L., Zhu, H., Tan, C.: A Survey of Embodied AI: From Simulators to Research Tasks. IEEE Transactions on Emerging Topics in Compu- tational Intelligence6, 230–244 (2022)

  21. [29]

    Efron, R.: What is Perception? In: Proceedings of the Boston Colloquium for the Philosophy of Science 1966/1968. pp. 137–173 (1969)

  22. [30]

    In: European Conference on Computer Vision

    Fan, Z., Ohkawa, T., Yang, L., Lin, N., Zhou, Z., Zhou, S., Liang, J., Gao, Z., Zhang, X., Zhang, X., et al.: Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects. In: European Conference on Computer Vision. pp. 428–448. Springer (2024)

  23. [31]

    In: European Conference on Computer Vision

    Farkhondeh, A., Tafasca, S., Odobez, J.M.: ChildPlay-Hand: A Dataset of Hand Manipulations in the Wild. In: European Conference on Computer Vision. pp. 101–118. Springer (2024)

  24. [32]

    Studia Linguistica78(1), 1–7 (2024)

    Fortuny, J., Payrató, L.: Ambiguity in Linguistics. Studia Linguistica78(1), 1–7 (2024)

  25. [33]

    arXiv preprint arXiv:2602.17573 (2026) 18 L

    Foteinos, K., Angelidis, G., Psiris, A., Argyriou, V., Sarigiannidis, P., Papadopou- los, G.T.: FR-GESTURE: An RGBD Dataset For Gesture-based Human-Robot Interaction In First Responder Operations. arXiv preprint arXiv:2602.17573 (2026) 18 L. Li et al

  26. [34]

    ArXiv Preprint ArXiv:2504.11346 (2025)

    Gao, Y., Gong, L., Guo, Q., Hou, X., Lai, Z., Li, F., Li, L., Lian, X., Liao, C., Liu, L., et al.: Seedream 3.0 Technical Report. ArXiv Preprint ArXiv:2504.11346 (2025)

  27. [35]

    IEEE Robotics & Automation Magazine14, 90–103 (2007)

    Garcia, E., Jimenez, M.A., De Santos, P.G., Armada, M.: The Evolution of Robotics Research. IEEE Robotics & Automation Magazine14, 90–103 (2007)

  28. [36]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., et al.: Ego4D: Around the World in 3,000 Hours of Egocentric Video. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp...

  29. [37]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Grauman, K., Westbury, A., Torresani, L., Kitani, K., Malik, J., Afouras, T., Ashutosh, K., Baiyya, V., Bansal, S., Boote, B., et al.: Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives. In: Proceedings of the IEEE/CVF Conference on Compu...

  30. [38]

    5-vl Technical Report

    Guo, D., Wu, F., Zhu, F., Leng, F., Shi, G., Chen, H., Fan, H., Wang, J., Jiang, J., Wang, J., et al.: Seed1. 5-vl Technical Report. ArXiv Preprint ArXiv:2505.07062 (2025)

  31. [39]

    ArXiv Preprint ArXiv:2401.17766 (2024)

    Guo, J., Rao, Z., Chen, Z., Zhou, J., Tao, D.: Fine-Grained Zero-Shot Learning: Advances, Challenges, and Prospects. ArXiv Preprint ArXiv:2401.17766 (2024)

  32. [41]

    IEEE Transactions on Human-Machine Systems51, 300–309 (2021)

    Guo, L., Lu, Z., Yao, L.: Human-Machine Interaction Sensing Technology Based on Hand Gesture Recognition: A Review. IEEE Transactions on Human-Machine Systems51, 300–309 (2021)

  33. [42]

    IEEE Access12, 143599–143626 (2024)

    Hashi, A.O., Hashim, S.Z.M., Asamah, A.B.: A Systematic Review of Hand Gesture Recognition: An Update From 2018 to 2024. IEEE Access12, 143599–143626 (2024)

  34. [43]

    In: Proceedings of the IEEE International Conference on Computer Vision

    He, K., Gkioxari, G., Dollár, P., Girshick, R.: Mask R-CNN. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 2961–2969 (2017)

  35. [44]

    In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition

    He, K., Zhang, X., Ren, S., Sun, J.: Deep Residual Learning for Image Recogni- tion. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 770–778 (2016)

  36. [45]

    In: European Conference on Computer Vision

    Hong Da, N.L., Miura, J., Hayashi, K.: Enhancing Human-Robot Collaborative Search Through Efficient Space Sharing with On-Demand Bidirectional Interaction. In: European Conference on Computer Vision. pp. 20–35. Springer (2024)

  37. [46]

    In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition

    Hu, R., Rohrbach, M., Andreas, J., Darrell, T., Saenko, K.: Modeling Relationships in Referential Expressions With Compositional Modular Networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 1115– 1124 (2017)

  38. [47]

    In: European Conference on Computer Vision

    Huang, Z., Satoh, S.: LoA-Trans: Enhancing Visual Grounding by Location-Aware Transformers. In: European Conference on Computer Vision. pp. 405–421. Springer (2024)

  39. [48]

    In: European Conference on Computer Vision

    Hummel, T., Karthik, S., Georgescu, M.I., Akata, Z.: EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval. In: European Conference on Computer Vision. pp. 1–17. Springer (2024)

  40. [49]

    Machines 11, 677 (2023)

    Hussain, M.: YOLO-v1 to YOLO-v8, the Rise of YOLO and Its Complementary Nature toward Digital Manufacturing and Industrial Defect Detection. Machines 11, 677 (2023)

  41. [50]

    Cities152, 1–94 (2024) Abbreviated paper title 19

    Ito, K., Kang, Y., Zhang, Y., Zhang, F., Biljecki, F.: Understanding Urban Perception with Visual Data: A Systematic Review. Cities152, 1–94 (2024) Abbreviated paper title 19

  42. [51]

    In: European Conference on Computer Vision

    Jiang, H., Lu, Z.: Visual Grounding for Object-Level Generalization in Reinforce- ment Learning. In: European Conference on Computer Vision. pp. 55–72. Springer (2024)

  43. [52]

    Neural Computing and Applications33(7), 2297–2319 (2021)

    Jirak, D., Biertimpel, D., Kerzel, M., Wermter, S.: Solving Visual Object Ambigu- ities when Pointing: An Unsupervised Learning Approach. Neural Computing and Applications33(7), 2297–2319 (2021)

  44. [53]

    In: European Conference on Computer Vision

    Ju, Y., Hu, K., Zhang, G., Zhang, G., Jiang, M., Xu, H.: Robo-ABC: Affordance Generalization Beyond Categories via Semantic Correspondence for Robot Ma- nipulation. In: European Conference on Computer Vision. pp. 222–239. Springer (2024)

  45. [54]

    In: Proceedings of the Computer Vision and Pattern Recognition Conference

    Kang, S., Kim, J., Kim, J., Hwang, S.J.: Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 9339–9350 (2025)

  46. [55]

    In: European Conference on Computer Vision

    Kang, W., Liu, G., Shah, M., Yan, Y.: SegVG: Transferring Object Bounding Box to Segmentation for Visual Grounding. In: European Conference on Computer Vision. pp. 57–75. Springer (2024)

  47. [56]

    In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing

    Kazemzadeh, S., Ordonez, V., Matten, M., Berg, T.: ReferItGame: Referring to Objects in Photographs of Natural Scenes. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing. pp. 787–798 (2014)

  48. [57]

    In: 2024 International Conference on Information Technology Research and Innovation (ICITRI)

    Kusnanti, E.A., Sierra, E., Putra, G.G.S., Cahyadi, E.S., Haq, A., Purwitasari, D.: Indonesian Lexical Ambiguity in Machine Translation: A Literature Review. In: 2024 International Conference on Information Technology Research and Innovation (ICITRI). pp. 59–64. IEEE (2024)

  49. [58]

    In: Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems

    Lee, J., Wang, J., Brown, E., Chu, L., Rodriguez, S.S., Froehlich, J.E.: Gaze- PointAR: A context-aware multimodal voice assistant for pronoun disambiguation in wearable augmented reality. In: Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. pp. 1–...

  50. [59]

    In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics

    Li, Z., Xu, Q., Zhang, D., Song, H., Cai, Y., Qi, Q., Zhou, R., Pan, J., Li, Z., Tu, V., et al.: GroundingGPT: Language Enhanced Multi-modal Grounding Model. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. pp. 6657–6678 (2024)

  51. [60]

    In: Proceedings of the European Conference on Computer Vision

    Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft COCO: Common Objects in Context. In: Proceedings of the European Conference on Computer Vision. pp. 740–755 (2014)

  52. [61]

    ArXiv Preprint ArXiv:1712.09005 (2017)

    Linderman, G.C., Rachh, M., Hoskins, J.G., Steinerberger, S., Kluger, Y.: Efficient Algorithms for t-distributed Stochastic Neighborhood Embedding. ArXiv Preprint ArXiv:1712.09005 (2017)

  53. [62]

    ArXiv Preprint ArXiv:2412.19437 (2024)

    Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., et al.: DeepSeek-V3 Technical Report. ArXiv Preprint ArXiv:2412.19437 (2024)

  54. [63]

    ArXiv Preprint ArXiv:2405.14295 (2024)

    Liu, C., Wei, H., Chen, J., Kong, L., Ge, Z., Zhu, Z., Zhao, L., Sun, J., Han, C., Zhang, X.: Focus Anywhere for Fine-grained Multi-page Document Understanding. ArXiv Preprint ArXiv:2405.14295 (2024)

  55. [64]

    In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

    Liu, H., Li, C., Li, Y., Lee, Y.J.: Improved baselines with visual instruction tuning. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 26296–26306 (2024)

  56. [65]

    ArXiv Preprint ArXiv:2405.07801 (2024) 20 L

    Liu, J., Sun, W., Yang, H., Zeng, Z., Liu, C., Zheng, J., Liu, X., Rahmani, H., Sebe, N., Mian, A.: Deep Learning-Based Object Pose Estimation: A Comprehensive Survey. ArXiv Preprint ArXiv:2405.07801 (2024) 20 L. Li et al

  57. [66]

    In: Proceedings of the IEEE/CVF International Conference on Computer Vision

    Liu, J., Chen, Q., Wang, Z., Tang, Y., Zhang, Y., Yan, C., Wang, D., Li, X., Zhao, B.: AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 5177–5187 (2025)

  58. [67]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Liu, R., Liu, C., Bai, Y., Yuille, A.L.: CLEVR-Ref+: Diagnosing Visual Reasoning With Referring Expressions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 4185–4194 (2019)

  59. [68]

    In: Proceedings of the European Conference on Computer Vision

    Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Jiang, Q., Li, C., Yang, J., Su, H., et al.: Grounding DINO: Marrying DINO with Grounded Pre-training for Open-Set Object Detection. In: Proceedings of the European Conference on Computer Vision. pp. 38–55 (2024)

  60. [69]

    IEEE/ASME Transactions on Mechatronics pp

    Liu,Y.,Chen,W.,Bai,Y.,Liang,X.,Li,G.,Gao,W.,Lin,L.:AligningCyberSpace With Physical World: A Comprehensive Survey on Embodied AI. IEEE/ASME Transactions on Mechatronics pp. 1–22 (2025)

  61. [70]

    International Journal of Robotics & Control Systems2(4) (2022)

    Ma’arif, A., Rahmaniar, W., Fathurrahman, H.I.K., Frisky, A.Z.K., et al.: Un- derstanding of Convolutional Neural Network (CNN): A Review. International Journal of Robotics & Control Systems2(4) (2022)

  62. [71]

    In: European Conference on Computer Vision

    Magay, A., Tripathi, D., Hao, Y., Fang, Y.: A Light and Smart Wearable Platform with Multimodal Foundation Model for Enhanced Spatial Reasoning in People with Blindness and Low Vision. In: European Conference on Computer Vision. pp. 323–339. Springer (2024)

  63. [72]

    In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition

    Mao, J., Huang, J., Toshev, A., Camburu, O., Yuille, A.L., Murphy, K.: Generation and Comprehension of Unambiguous Object Descriptions. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 11–20 (2016)

  64. [73]

    In: European Conference on Computer Vision

    Memmi, M.Y., Slama, R., Berretti, S.: Hand Gesture Recognition Using Dual Graph Hierarchical Edges Representation and Graph Transformer Network. In: European Conference on Computer Vision. pp. 53–68. Springer (2024)

  65. [74]

    Mo, M., Wang, J., Wang, L., Chen, H., Gu, C., Leng, J., Gao, X.: NexusAD: Exploring the Nexus for Multimodal Perception and Comprehension of Corner CasesinAutonomousDriving.In:ECCV2024WorkshoponMultimodalPerception and Comprehension of Corner Cases in Autonomous Driving (2024)

  66. [75]

    In: Proceedings of the European Conference on Computer Vision

    Nagaraja, V.K., Morariu, V.I., Davis, L.S.: Modeling Context between Objects for Referring Expression Understanding. In: Proceedings of the European Conference on Computer Vision. pp. 792–807 (2016)

  67. [76]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Narasimhaswamy, S., Bhattacharya, U., Chen, X., Dasgupta, I., Mitra, S., Hoai, M.: HanDiffuser: Text-to-Image Generation With Realistic Hand Appearances. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 2468–2479 (2024)

  68. [77]

    ArXiv Preprint ArXiv:2304.07193 (2023)

    Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al.: DINOv2: Learning Robust Visual Features without Supervision. ArXiv Preprint ArXiv:2304.07193 (2023)

  69. [78]

    ArXiv Preprint ArXiv:2306.14824 (2023)

    Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., Wei, F.: Kosmos-2: Grounding Multimodal Large Language Models to the World. ArXiv Preprint ArXiv:2306.14824 (2023)

  70. [79]

    In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition

    Pepikj, B., Stark, M., Gehler, P., Schiele, B.: Occlusion Patterns for Object Class Detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 3286–3293 (2013)

  71. [80]

    Lecture Notes in Computer Science pp

    Pfeifer, R., Iida, F.: Embodied Artificial Intelligence: Trends and Challenges. Lecture Notes in Computer Science pp. 1–26 (2004) Abbreviated paper title 21

  72. [81]

    In: Proceedings of the IEEE International Conference on Computer Vision

    Plummer, B.A., Wang, L., Cervantes, C.M., Caicedo, J.C., Hockenmaier, J., Lazebnik, S.: Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 2641–2649 (2015)

  73. [82]

    IEEE Transactions on Pattern Analysis and Machine Intelligence45, 4051–4070 (2022)

    Pourpanah,F.,Abdar,M.,Luo,Y.,Zhou,X.,Wang,R.,Lim,C.P.,Wang,X.Z.,Wu, Q.J.: A Review of Generalized Zero-Shot Learning Methods. IEEE Transactions on Pattern Analysis and Machine Intelligence45, 4051–4070 (2022)

  74. [83]

    In: European Conference on Computer Vision

    Qi, Z., Dong, R., Zhang, S., Geng, H., Han, C., Ge, Z., Yi, L., Ma, K.: ShapeLLM: Universal 3D Object Understanding for Embodied Interaction. In: European Conference on Computer Vision. pp. 214–238. Springer (2024)

  75. [84]

    IEEE Transactions on Multimedia23, 4426–4440 (2020)

    Qiao, Y., Deng, C., Wu, Q.: Referring Expression Comprehension: A Survey of Methods and Datasets. IEEE Transactions on Multimedia23, 4426–4440 (2020)

  76. [85]

    In: Proceedings of the International Conference on Machine Learning

    Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning Transferable Visual Models From Natural Language Supervision. In: Proceedings of the International Conference on Machine Learning. pp. 8748–8...

  77. [86]

    In: Proceedings of the International Conference on Multimodal Interfaces and the Workshop on Machine Learning for Multimodal Interaction

    Schauerte, B., Fink, G.A.: Focusing Computational Visual Attention in Multi- Modal Human-Robot Interaction. In: Proceedings of the International Conference on Multimodal Interfaces and the Workshop on Machine Learning for Multimodal Interaction. pp. 1–8 (2010)

  78. [87]

    In: Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems

    Schauerte, B., Richarz, J., Fink, G.A.: Saliency-Based Identification and Recog- nition of Pointed-At Objects. In: Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems. pp. 4638–4643 (2010)

  79. [88]

    In: European Conference on Computer Vision

    Shi, C., Yang, S.: Spatial and Visual Perspective-Taking via View Rotation and Relation Reasoning for Embodied Reference Understanding. In: European Conference on Computer Vision. pp. 201–218. Springer (2022)

  80. [89]

    In: Proceedings of the 2015 International Conference on Digital Image Computing: Techniques and Applications

    Shukla, D., Erkent, O., Piater, J.: Probabilistic Detection of Pointing Directions for Human-Robot Interaction. In: Proceedings of the 2015 International Conference on Digital Image Computing: Techniques and Applications. pp. 1–8 (2015)

  81. [90]

    In: Proceedings of the IEEE International Symposium on Robot and Human Interactive Communication

    Shukla, D., Erkent, Ö., Piater, J.: A Multi-View Hand Gesture RGB-D Dataset for Human-Robot Interaction Scenarios. In: Proceedings of the IEEE International Symposium on Robot and Human Interactive Communication. pp. 1084–1091 (2016)

  82. [91]

    Su, W., Miao, P., Dou, H., Li, X.: Scanformer: Referring Expression Comprehension byIterativelyScanning.In:ProceedingsoftheIEEE/CVFConferenceonComputer Vision and Pattern Recognition. pp. 13449–13458 (2024)

  83. [92]

    arXiv preprint arXiv:2312.11805 (2023)

    Team, G., Anil, R., Borgeaud, S., Alayrac, J.B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A.M., Hauth, A., Millican, K., et al.: Gemini: A Family of Highly Capable Multimodal Models. arXiv preprint arXiv:2312.11805 (2023)

  84. [93]

    ArXiv Preprint ArXiv:2503.20020 (2025)

    Team, G.R., Abeyruwan, S., Ainslie, J., Alayrac, J.B., Arenas, M.G., Armstrong, T., Balakrishna, A., Baruch, R., Bauza, M., Blokzijl, M., et al.: Gemini Robotics: Bringing AI into the Physical World. ArXiv Preprint ArXiv:2503.20020 (2025)

  85. [94]

    ArXiv Preprint ArXiv:2302.13971 (2023)

    Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al.: LLaMA: Open and Efficient Foundation Language Models. ArXiv Preprint ArXiv:2302.13971 (2023)

  86. [95]

    In: Proceedings of the Advances in Neural Information Processing Systems

    Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is All you Need. In: Proceedings of the Advances in Neural Information Processing Systems. pp. 1–11 (2017)

  87. [96]

    Eye38, 1036–1038 (2024) 22 L

    Waisberg, E., Ong, J., Masalkhi, M., Zaman, N., Sarker, P., Lee, A.G., Tavakkoli, A.: Meta Smart Glasses—Large Language Models and the Future for Assistive Glasses for Individuals with Vision Impairments. Eye38, 1036–1038 (2024) 22 L. Li et al

  88. [97]

    ArXiv Preprint ArXiv:2503.07465 (2025)

    Wang, A., Liu, L., Chen, H., Lin, Z., Han, J., Ding, G.: YOLOE: Real-Time Seeing Anything. ArXiv Preprint ArXiv:2503.07465 (2025)

  89. [98]

    ArXiv Preprint ArXiv:2305.16291 (2023)

    Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., Anandku- mar, A.: Voyager: An Open-Ended Embodied Agent with Large Language Models. ArXiv Preprint ArXiv:2305.16291 (2023)

  90. [99]

    Journal of Automation and Intelligence4, 52–64 (2025)

    Wang, J., Shi, E., Hu, H., Ma, C., Liu, Y., Wang, X., Yao, Y., Liu, X., Ge, B., Zhang, S.: Large Language Models for Robotics: Opportunities, Challenges, and Perspectives. Journal of Automation and Intelligence4, 52–64 (2025)

  91. [100]

    ArXiv Preprint ArXiv:2409.12191 (2024)

    Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., et al.: Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution. ArXiv Preprint ArXiv:2409.12191 (2024)

  92. [101]

    ArXiv Preprint ArXiv:2508.18265 (2025)

    Wang, W., Gao, Z., Gu, L., Pu, H., Cui, L., Wei, X., Liu, Z., Jing, L., Ye, S., Shao, J., et al.: InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency. ArXiv Preprint ArXiv:2508.18265 (2025)

  93. [102]

    Journal of Vision2(7), 435–435 (2002)

    Watanabe, K.: Reflexive Attentional Shift Caused by Indexical Pointing Gesture. Journal of Vision2(7), 435–435 (2002)

  94. [103]

    arXiv preprint arXiv:2412.10302 (2024)

    Wu, Z., Chen, X., Pan, Z., Liu, X., Liu, W., Dai, D., Gao, H., Ma, Y., Wu, C., Wang, B., et al.: DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding. arXiv preprint arXiv:2412.10302 (2024)

  95. [104]

    In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition

    Xian, Y., Schiele, B., Akata, Z.: Zero-Shot Learning - the Good, the Bad and the Ugly. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 4582–4591 (2017)

  96. [105]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Xiao, B., Wu, H., Xu, W., Dai, X., Hu, H., Lu, Y., Zeng, M., Liu, C., Yuan, L.: Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 4818–4829 (2024)

  97. [106]

    ArXiv Preprint ArXiv:2412.20206 (2024)

    Xiao, L., Yang, X., Lan, X., Wang, Y., Xu, C.: Towards Visual Grounding: A Survey. ArXiv Preprint ArXiv:2412.20206 (2024)

  98. [107]

    IEEE Transactions on Pattern Analysis and Machine Intelligence (2025)

    Xiao, L., Yang, X., Lan, X., Wang, Y., Xu, C.: Towards Visual Grounding: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence (2025)

  99. [108]

    In: ProceedingsoftheAdvancesinNeuralInformationProcessingSystems.pp.139854– 139885 (2024)

    Xiao, L., Yang, X., Peng, F., Wang, Y., Xu, C.: OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling. In: ProceedingsoftheAdvancesinNeuralInformationProcessingSystems.pp.139854– 139885 (2024)

  100. [109]

    In: European Conference on Computer Vision

    Xie, P., Chen, S., Dai, Y., Hu, D., Yang, K., Wang, G.: Target-Oriented Object Grasping via Multimodal Human Guidance. In: European Conference on Computer Vision. pp. 305–322. Springer (2024)

  101. [110]

    ArXiv Preprint ArXiv:2505.09388 (2025)

    Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., et al.: Qwen3 Technical Report. ArXiv Preprint ArXiv:2505.09388 (2025)

  102. [111]

    In: European conference on computer vision

    Yang,J.,Ding,R.,Brown,E.,Qi,X.,Xie,S.:V-IRL:GroundingVirtualIntelligence in Real Life. In: European conference on computer vision. pp. 36–55. Springer (2024)

  103. [112]

    In: Proceedings of the IEEE/CVF International Conference on Computer Vision

    Yang, S., Li, G., Yu, Y.: Dynamic Graph Attention for Referring Expression Comprehension. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 4644–4653 (2019)

  104. [113]

    ArXiv Preprint ArXiv:2310.07704 (2023) Abbreviated paper title 23

    You, H., Zhang, H., Gan, Z., Du, X., Zhang, B., Wang, Z., Cao, L., Chang, S.F., Yang, Y.: Ferret: Refer and Ground Anything Anywhere at Any Granularity. ArXiv Preprint ArXiv:2310.07704 (2023) Abbreviated paper title 23

  105. [114]

    In: Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition

    Yu, L., Lin, Z., Shen, X., Yang, J., Lu, X., Bansal, M., Berg, T.L.: MAttNet: Modular Attention Network for Referring Expression Comprehension. In: Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 1307–1315 (2018)

  106. [115]

    In: Proceedings of the European Conference on Computer Vision

    Yu, L., Poirson, P., Yang, S., Berg, A.C., Berg, T.L.: Modeling Context in Referring Expressions. In: Proceedings of the European Conference on Computer Vision. pp. 69–85 (2016)

  107. [116]

    In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition

    Zhang, H., Niu, Y., Chang, S.F.: Grounding Referring Expressions in Images by Variational Context. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 4158–4166 (2018)

  108. [117]

    ArXiv Preprint ArXiv:2203.03605 (2022)

    Zhang, H., Li, F., Liu, S., Zhang, L., Su, H., Zhu, J., Ni, L.M., Shum, H.Y.: DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection. ArXiv Preprint ArXiv:2203.03605 (2022)

  109. [118]

    In: Proceedings of the European Conference on Computer Vision

    Zhang, H., Li, H., Li, F., Ren, T., Zou, X., Liu, S., Huang, S., Gao, J., Leizhang, Li, C., et al.: LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models. In: Proceedings of the European Conference on Computer Vision. pp. 19–35 (2024)

  110. [119]

    Visual Informatics6, 56–67 (2022)

    Zhao, Y., Jiang, J., Chen, Y., Liu, R., Yang, Y., Xue, X., Chen, S.: Metaverse: Perspectives from Graphics, Interactions and Visualization. Visual Informatics6, 56–67 (2022)

  111. [120]

    what is happening,

    Zhuang, B., Wu, Q., Shen, C., Reid, I., Van Den Hengel, A.: Parallel Attention: A Unified Framework for Visual Object Discovery Through Dialogs and Queries. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 4252–4261 (2018) 24 L. Li et al. ...

Pith tools

Reviewed July 14, 2026 · model on record in the stance chip above.