REVIEW 2 major objections 1 minor 45 references
SuperFashion replaces patch-based attention with superpixel tokens in a Transformer to improve attribute-specific fashion retrieval.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-27 11:39 UTC pith:KNYD2HLS
load-bearing objection Superpixel tokens guided by attribute attention are positioned as the key change for ASFR, but the reported MAP gains are not isolated from other modeling choices. the 2 major comments →
Beyond Patches: Superpixel Token-based Transformers for Attribute-Specific Fashion Retrieval
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
SuperFashion is the first ASFR framework that adopts superpixel tokens within a Transformer architecture. It employs an attribute-guided attention mechanism to extract attribute-related features and guide cropping of semantically meaningful regions, then uses superpixel segmentation on those regions to generate compact, semantically coherent superpixel tokens. By incorporating modality-specific embeddings for both attribute and superpixel tokens, the superpixel token-based Transformer facilitates adaptive interaction and fusion, enhancing attribute localization and discrimination, with relative overall MAP improvements of 1.84%, 9.27%, and 9.35% over prior SOTA on FashionAI, DARN, and DeepFa
What carries the argument
The superpixel token-based Transformer, which combines attribute-guided attention for region cropping with superpixel segmentation to produce tokens that enable adaptive interaction and fusion.
Load-bearing premise
Attribute-guided attention reliably extracts attribute-related features for cropping, and superpixel segmentation on those regions yields tokens that capture pixel-level microstructures better than patches.
What would settle it
Running the same retrieval experiments on FashionAI, DARN, and DeepFashion but replacing superpixel segmentation with standard patch division and finding no MAP improvement or a decline would falsify the central advantage.
If this is right
- The method reduces background noise and misalignment with irregular attribute regions compared with patch-based Transformers.
- Attribute localization and discrimination improve through adaptive fusion of attribute and superpixel tokens.
- Web-based image retrieval systems can achieve higher precision on fine-grained fashion queries.
- Superpixel tokens offer a more semantically coherent alternative to fixed patches for microstructure capture.
Where Pith is reading between the lines
- The cropping-plus-segmentation pipeline could be tested on non-fashion domains that also feature irregular target regions.
- Fewer superpixel tokens than patches might lower memory use during Transformer inference, though the paper does not measure this.
- The same token construction could be paired with other attention mechanisms beyond the attribute-guided one used here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces SuperFashion, the first ASFR framework to adopt superpixel tokens within a Transformer architecture. It employs an attribute-guided attention mechanism to extract relevant features and crop semantically meaningful regions, applies superpixel segmentation to generate compact tokens, and incorporates modality-specific embeddings to enable adaptive interaction and fusion between attribute and superpixel tokens. Experiments on FashionAI, DARN, and DeepFashion report relative overall MAP improvements of 1.84%, 9.27%, and 9.35% over prior SOTA.
Significance. If the performance gains can be robustly attributed to the superpixel token mechanism, the work would represent a targeted advance in fine-grained fashion retrieval by addressing misalignment with irregular attribute regions and background noise through semantically coherent tokens rather than fixed patches. This could influence subsequent designs for attribute-specific tasks in image retrieval.
major comments (2)
- [Experiments] The central claim that superpixel tokens (generated after attribute-guided cropping) drive the reported MAP gains rests on an untested assumption; no ablation studies or token-size comparisons isolate this component from modality-specific embeddings, Transformer depth, or the attention mechanism itself (Experiments section).
- [Abstract and Experiments] Performance numbers in the abstract and Experiments section are presented without baselines, dataset splits, error bars, or statistical tests, preventing verification that the 1.84–9.35% relative improvements are reliable or attributable to the proposed method.
minor comments (1)
- [Abstract] The final sentence of the abstract ('SuperFashion offers a new solution for web-based image retrieval.') is vague and could be removed or made more specific to the contribution.
Simulated Author's Rebuttal
We thank the referee for the insightful comments on our work. We address each major comment below and outline the revisions we will make to strengthen the manuscript.
read point-by-point responses
-
Referee: [Experiments] The central claim that superpixel tokens (generated after attribute-guided cropping) drive the reported MAP gains rests on an untested assumption; no ablation studies or token-size comparisons isolate this component from modality-specific embeddings, Transformer depth, or the attention mechanism itself (Experiments section).
Authors: We agree that the manuscript would benefit from ablation studies to isolate the contribution of the superpixel tokens. In the revised version, we will add comprehensive ablation experiments, including comparisons with and without superpixel tokens, different token sizes, and controls for modality-specific embeddings, Transformer depth, and the attention mechanism. These will be presented in the Experiments section to better support our central claim. revision: yes
-
Referee: [Abstract and Experiments] Performance numbers in the abstract and Experiments section are presented without baselines, dataset splits, error bars, or statistical tests, preventing verification that the 1.84–9.35% relative improvements are reliable or attributable to the proposed method.
Authors: We acknowledge the need for more detailed reporting to allow verification of the results. We will revise the abstract and Experiments section to include explicit baseline methods, details on dataset splits, error bars from multiple experimental runs, and results of statistical significance tests for the reported MAP improvements. revision: yes
Circularity Check
No circularity in proposed SuperFashion architecture
full rationale
The manuscript describes an empirical ML framework that combines attribute-guided attention for region cropping with superpixel segmentation to produce tokens for a Transformer. No equations, fitted parameters, or predictions are shown that reduce to the inputs by construction. No self-citations are invoked as load-bearing uniqueness theorems or ansatzes. Performance claims rest on reported MAP deltas versus prior SOTA on public datasets rather than any definitional or self-referential reduction. The derivation chain is therefore self-contained and non-circular.
Axiom & Free-Parameter Ledger
read the original abstract
Attribute-Specific Fashion Retrieval (ASFR) aims to improve fine-grained image retrieval by focusing on specific attributes. However, existing patch-based attention and Transformer methods often misalign with irregular attribute regions and are prone to background noise, limiting their ability to capture subtle, pixel-level microstructures. To tackle these challenges, we propose SuperFashion, the first ASFR framework that adopts superpixel tokens within a Transformer architecture. SuperFashion initially employs an attribute-guided attention mechanism to extract attribute-related features, which in turn guide the cropping of semantically meaningful image regions. Superpixel segmentation is then leveraged on these regions to generate compact, semantically coherent superpixel tokens. By incorporating modality-specific embeddings for both attribute and superpixel tokens, the superpixel token-based Transformer facilitates adaptive interaction and fusion, thereby enhancing attribute localization and discrimination. Extensive experiments on FashionAI, DARN, and DeepFashion demonstrate relative overall MAP improvements of 1.84%, 9.27%, and 9.35% over prior SOTA. SuperFashion offers a new solution for web-based image retrieval.
Figures
Reference graph
Works this paper leans on
-
[1]
Marius Aasan, Odd Kolbjørnsen, Anne Schistad Solberg, and Adín Ramírez Rivera
-
[2]
InEuropean Conference on Computer Vision
A Spitting Image: Modular Superpixel Tokenization in Vision Transformers. InEuropean Conference on Computer Vision. Springer, Cham, Switzerland, 124– 142
-
[3]
Radhakrishna Achanta, Appu Shaji, Kevin Smith, Aurelien Lucchi, Pascal Fua, and Sabine Süsstrunk. 2012. SLIC Superpixels Compared to State-of-the-Art Su- perpixel Methods.IEEE Transactions on Pattern Analysis and Machine Intelligence 34, 11 (2012), 2274–2282
2012
-
[4]
Isabela Borlido Barcelos, Felipe De Castro Belém, Leonardo De Melo João, Ze- nilton KG Do Patrocínio Jr, Alexandre Xavier Falcão, and Silvio Jamil Ferzoli Guimarães. 2024. A Comprehensive Review and New Taxonomy on Superpixel Segmentation.Comput. Surveys56, 8 (2024), 1–39
2024
-
[5]
Antonio D’Innocente, Nikhil Garg, Yuan Zhang, Loris Bazzani, and Michael Donoser. 2021. Localized Triplet Loss for Fine-Grained Fashion Image Retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3910–3915
2021
-
[6]
Jianfeng Dong, Zhe Ma, Xiaofeng Mao, Xun Yang, Yuan He, Richang Hong, and Shouling Ji. 2021. Fine-Grained Fashion Similarity Prediction by Attribute- Specific Embedding Learning.IEEE Transactions on Image Processing30 (2021), 8410–8425
2021
-
[7]
Jianfeng Dong, Xiaoman Peng, Zhe Ma, Daizong Liu, Xiaoye Qu, Xun Yang, Jixiang Zhu, and Baolong Liu. 2023. From Region to Patch: Attribute-Aware Foreground-Background Contrastive Learning for Fine-Grained Fashion Re- trieval. InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1273–1282
2023
-
[8]
Garas Gendy, Guanghui He, and Nabil Sabor. 2023. Lightweight Image Super- Resolution Based on Deep Learning: State-of-the-Art and Future Directions. Information Fusion94 (2023), 284–310
2023
-
[9]
Kankanhalli
Xiaoling Gu, Yongkang Wong, Lidan Shou, Pai Peng, Gang Chen, and Mohan S. Kankanhalli. 2019. Multi-Modal and Multi-Domain Embedding Learning for Fashion Retrieval and Analysis.IEEE Transactions on Multimedia21, 6 (2019), 1524–1537
2019
-
[10]
Yunpeng Han, Lisai Zhang, Qingcai Chen, Zhijian Chen, Zhonghua Li, Jianxin Yang, and Zhao Cao. 2023. FashionSAP: Symbols and Attributes Prompt for Fine- Grained Fashion Vision-Language Pre-Training. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 15028–15038
2023
-
[11]
Philippe Hansen-Estruch, David Yan, Ching-Yao Chuang, Orr Zohar, Jialiang Wang, Tingbo Hou, Tao Xu, Sriram Vishwanath, Peter Vajda, and Xinlei Chen
-
[12]
InProceedings of the 42nd International Conference on Machine Learning
Learnings from Scaling Visual Tokenizers for Reconstruction and Genera- tion. InProceedings of the 42nd International Conference on Machine Learning
-
[13]
Jakob Drachmann Havtorn, Amélie Royer, Tijmen Blankevoort, and Babak Eht- eshami Bejnordi. 2023. MSViT: Dynamic Mixed-scale Tokenization for Vision Transformers. InProceedings of the IEEE/CVF International Conference on Com- puter Vision. 838–848
2023
-
[14]
Feris, Qiang Chen, and Shuicheng Yan
Junshi Huang, Rogerio S. Feris, Qiang Chen, and Shuicheng Yan. 2015. Cross- Domain Image Retrieval with a Dual Attribute-Aware Ranking Network. In Proceedings of the IEEE International Conference on Computer Vision. 1062–1070
2015
-
[15]
Yang Jiao, Yan Gao, Jingjing Meng, Jin Shang, and Yi Sun. 2023. Learning Attribute and Class-Specific Representation Duet for Fine-Grained Fashion Analysis. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11050– 11059
2023
-
[16]
Yang Jiao, Ning Xie, Yan Gao, Chien-chih Wang, and Yi Sun. 2022. Fine-Grained Fashion Representation Learning by Online Deep Clustering. InUropean Confer- ence on Computer Vision. 19–35
2022
-
[17]
Sangtae Kim, Daeyoung Park, and Byonghyo Shim. 2023. Semantic-Aware Super- pixel for Weakly Supervised Semantic Segmentation. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 1142–1150
2023
-
[18]
Suha Kwak, Seunghoon Hong, and Bohyung Han. 2017. Weakly Supervised Semantic Segmentation Using Superpixel Pooling Network. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 31. 4111–4117
2017
- [19]
-
[20]
An-An Liu, Ting Zhang, Dan Song, Wenhui Li, and Ming Zhou. 2021. FRSFN: A Semantic Fusion Network for Practical Fashion Retrieval.Multimedia Tools and Applications80 (2021), 17169–17181
2021
-
[21]
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. InProceedings of the IEEE/CVF International Conference on Computer Vision. 10012–10022
2021
-
[22]
Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. 2016. DeepFash- ion: Powering Robust Clothes Recognition and Retrieval with Rich Annotations. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1096–1104
2016
-
[23]
Zhe Ma, Jianfeng Dong, Zhongzi Long, Yao Zhang, Yuan He, Hui Xue, and Shoul- ing Ji. 2020. Fine-Grained Fashion Similarity Learning by Attribute-Specific Embedding Network. InProceedings of the AAAI Conference on Artificial Intelli- gence. 11741–11748
2020
-
[24]
Tomer Ronen, Omer Levy, and Avram Golbert. 2023. Vision Transformers with Mixed-resolution Tokenization. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4613–4622
2023
-
[25]
Ronghua Shang, Jiyu Zhang, Licheng Jiao, Yangyang Li, Naresh Marturi, and Rustam Stolkin. 2020. Multi-Scale Adaptive Feature Fusion Network for Semantic Segmentation in Remote Sensing Images.Remote Sensing12, 5 (2020), 872
2020
-
[26]
Jianbing Shen, Xiaopeng Hao, Zhiyuan Liang, Yu Liu, Wenguan Wang, and Ling Shao. 2016. Real-time superpixel segmentation by DBSCAN clustering algorithm. IEEE Transactions on Image Processing25, 12 (2016), 5933–5942
2016
-
[27]
Chull Hwan Song and Hye Joo Han. 2022. Convolutional Attribute Mask with Two-Step Attention for Fashion Image Retrieval. InProceedings of the 2022 26th International Conference on Pattern Recognition. 2093–2099
2022
-
[28]
Matthew Tancik, Pratul Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T Barron, and Ren Ng. 2020. Fourier features let networks learn high frequency functions in low dimensional domains. InProceedings of the 34th Conference on Neural Information Processing Systems. 7537–7547
2020
-
[29]
Yuxin Tian, Shawn Newsam, and Kofi Boakye. 2023. Fashion Image Retrieval With Text Feedback by Additive Attention Compositional Learning. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 1011–1021
2023
-
[30]
Manmatha, and C
Son Tran, Ming Du, Sampath Chanda, R. Manmatha, and C. J. Taylor. 2019. Searching for Apparel Products from Images in the Wild. InProceedings of the KDD 2019 Workshop on AI for Fashion
2019
-
[31]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need. InAdvances in Neural Information Processing Systems, Vol. 30. 5998– 6008
2017
-
[32]
Andreas Veit, Serge Belongie, and Theofanis Karaletsos. 2017. Conditional Similar- ity Networks. In2017 IEEE Conference on Computer Vision and Pattern Recognition. 1781–1789
2017
-
[33]
Yongquan Wan, Kang Yan, Cairong Yan, and Bofeng Zhang. 2024. Learning Attribute-Guided Fashion Similarity with Spatial and Channel Attention.Journal of Experimental & Theoretical Artificial Intelligence36, 5 (2024), 703–719
2024
-
[34]
Ling Xiao and Toshihiko Yamasaki. 2024. Boosting Fine-grained Fashion Retrieval with Relational Knowledge Distillation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 8229–8234
2024
-
[35]
Ling Xiao and Toshihiko Yamasaki. 2025. GeoDCL: Weak Geometrical Distortion Based Contrastive Learning for Fine-Grained Fashion Image Retrieval.IEEE Transactions on Artificial Intelligence6, 3 (2025), 1234–1245
2025
-
[36]
Zhenwei Xie, Bing Wang, Zhanqiang Liu, Liping Jiang, and Yang Liu. 2025. A Novel Superpixel Segmentation Method Based on Adaptive Seed Expansion Random Walk Algorithm for Complex Scene Images.IEEE Transactions on Instrumentation and Measurement74 (2025), 1–12
2025
-
[37]
Cairong Yan, Anan Ding, Yanting Zhang, and Zijian Wang. 2021. Learning Fashion Similarity Based on Hierarchical Attribute Embedding. In2021 IEEE 8th International Conference on Data Science and Advanced Analytics. IEEE, 1–8
2021
-
[38]
Cairong Yan, Kang Yan, Yanting Zhang, Yongquan Wan, and Dandan Zhu. 2022. Attribute-Guided Fashion Image Retrieval by Iterative Similarity Learning. In 2022 IEEE International Conference on Multimedia and Expo. 1–6
2022
-
[39]
Tao Yang, Yuwang Wang, Yan Lu, and Nanning Zheng. 2022. Visual Concepts Tokenization. InAdvances in Neural Information Processing Systems, Vol. 35. 31571–31582
2022
-
[40]
Yue Yu, Yang Yang, and Kezhao Liu. 2021. Edge-Aware Superpixel Segmentation with Unsupervised Convolutional Neural Networks. InProceedings of the 2021 IEEE International Conference on Image Processing. 1504–1508
2021
-
[41]
Ross, Cordelia Schmid, Dina Katabi, and Xiuye Gu
Kaiwen Zha, Lijun Yu, Alireza Fathi, David A. Ross, Cordelia Schmid, Dina Katabi, and Xiuye Gu. 2025. Language-Guided Image Tokenization for Generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 15713–15722
2025
-
[42]
Shuili Zhang, Hongzhang Mu, Tingwen Liu, Qianqian Tong, and Jiawei Sheng
-
[43]
InProceedings of the 33rd ACM International Conference on Information and Knowledge Management, CIKM 2024, Boise, ID, USA, October 21-25, 2024
MSKR: Advancing Multi-modal Structured Knowledge Representation with Synergistic Hard Negative Samples. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management, CIKM 2024, Boise, ID, USA, October 21-25, 2024. ACM, 3207–3216
2024
-
[44]
Alex Zihao Zhu, Jieru Mei, Siyuan Qiao, Hang Yan, Yukun Zhu, Liang-Chieh Chen, and Henrik Kretzschmar. 2023. Superpixel Transformers for Efficient Semantic Segmentation. InProceedings of the 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems. 7651–7658. WWW ’26, April 13–17, 2026, Dubai, United Arab Emirates Shuili Zhang, Hongzhang M...
2023
-
[45]
Xingxing Zou, Xiangheng Kong, Waikeung Wong, Congde Wang, Yuguang Liu, and Yang Cao. 2019. FashionAI: A Hierarchical Dataset for Fashion Understand- ing. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops
2019
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.