REVIEW 4 major objections 3 minor 73 references
Symmetry Understanding of 3D Shapes via Chirality Disentanglement
T0 review · 4 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The paper claims that per-vertex chirality-aware features for 3D shapes can be extracted unsupervised from 2D foundation models via view-based lifting, so that descriptors can distinguish left from right symmetric parts while staying robust
desk verdict The supplied text is not the chirality paper; the abstract is unverifiable, so desk reject pending the correct PDF. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
An unsupervised, view-based lifting pipeline: for each vertex of the 3D shape, render the shape from many viewpoints, extract feature vectors from a frozen 2D foundation model at the projected pixel locations, and aggregate the per-view features into one per-vertex descriptor. The descriptor inherits the model's implicit sense of chirality and gives standard shape features a handedness channel.
What would settle it
Take a set of mirror-symmetric shapes with known left/right correspondences (e.g., left- and right-hand meshes), compute the per-vertex chirality features, and check whether mirror-matched vertices receive systematically different, stable signatures while ordinary rigidly aligned shapes do not. If the features are identical for mirror-image shapes or vary arbitrarily with the view set, the central claim fails.
Extended reading notes
Core claim
The central claim is that handedness information, which distinguishes a left hand from a right hand or the left and right sides of a body, can be recovered for 3D point clouds and meshes without any labels. Building on an existing view-lifting framework that paints 2D foundation-model features onto 3D vertices, the paper decorates each vertex with a chirality-aware feature describing its left/right identity. Because the 2D features come from models pretrained on images, the paper argues they encode cues that survive being projected back onto the 3D surface. The authors report that the resulting features, evaluated quantitatively and qualitatively across datasets, improve left-right disentang
Load-bearing premise
The load-bearing premise is that 2D foundation models retain enough handedness information that projecting it onto 3D vertices and aggregating across views preserves the left/right signal rather than washing it out.
Editorial extensions
If this is right
- Shape descriptors that were invariant to rigid motions can now tell mirror-symmetric points apart, so left/right disambiguation no longer needs manual correspondences or labels.
- Left-right disentanglement becomes a per-vertex property: objects like human bodies or animal shapes can be given a consistent left/right identity.
- Shape matching can use a chirality channel to avoid matching a point to its mirror counterpart.
- Part segmentation gains handedness as an additional cue to separate left and right parts of an object.
- Because extraction is unsupervised and uses frozen 2D models, the chirality channel adds no annotation cost to downstream tasks.
Reading between the lines
- The approach implicitly assumes a consistent global left/right reference can be defined across shapes; the abstract does not specify how this reference is anchored, making cross-shape consistency a testable open point.
- Many 2D foundation models are trained with horizontal-flip augmentation that deliberately removes left/right bias; whether chirality cues survive such training is the empirical crux, and could be tested by comparing features from models trained with and without flip augmentation.
- The full-text body supplied for this listing is an unrelated event-camera paper; the abstract is the only available evidence for the chirality pipeline's claims, so the datasets, metrics, and lifting details named in the abstract cannot be checked here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submitted manuscript, as identified by its title and abstract (arXiv:2508.05505), claims an unsupervised chirality feature extraction pipeline for 3D shapes. The abstract states that the method, built on the Diff3F framework, decorates shape vertices with chirality-aware information extracted from 2D foundation models, and that quantitative and qualitative experiments demonstrate effectiveness on left-right disentanglement, shape matching, and part segmentation. However, the full text supplied for review is a different paper: 'Revealing Latent Information: A Physics-inspired Self-supervised Pre-training Framework for Noisy and Sparse Events' by Lin Zhu et al. (arXiv:2508.05507v1, ACM MM '25). The body contains no description of chirality features, no equations or algorithms for the claimed pipeline, no mention of Diff3F, point clouds, meshes, or the listed downstream tasks, and no experimental results on left-right disentanglement, shape matching, or part segmentation. The submitted document therefore provides no support for the advertised central claim.
Significance. If the chirality method described in the abstract existed and were validated, it could address a real gap in 3D shape analysis: most rotation-invariant descriptors cannot distinguish mirror-symmetric parts, and a per-vertex chirality channel lifted from 2D foundation models would be a useful addition. However, because the manuscript body is an unrelated event-camera preprocessing paper, the significance of the claimed contribution cannot be assessed. The reviewer is not in a position to evaluate the correctness, novelty, or empirical strength of a method that is completely absent from the submitted text.
major comments (4)
- [Abstract vs. Full Text] The title and abstract propose a 'Symmetry Understanding of 3D Shapes via Chirality Disentanglement' method and promise downstream experiments on left-right disentanglement, shape matching, and part segmentation. The full text, however, is a different paper titled 'Revealing Latent Information: A Physics-inspired Self-supervised Pre-training Framework for Noisy and Sparse Events', with its own abstract, author list, ACM copyright block, and arXiv identifier (2508.05507v1). There is no substantive connection between the advertised claim and the body. This is a load-bearing evidentiary gap: the central claim of the manuscript cannot be checked against any method, equation, or result in the submitted document.
- [Method Description (missing)] The manuscript provides no description of the chirality feature extraction pipeline. There are no equations defining the per-vertex chirality channel, no specification of the view-based lifting from 2D foundation models, no aggregation rule, and no discussion of how parity information survives aggregation. Sections 1–5 and the supplementary material are exclusively about masked modeling, contrastive learning, event voxels, and downstream event-camera tasks. The claimed method is therefore entirely unverifiable from the supplied text.
- [Experiments and Results (missing)] The abstract claims that 'results from downstream tasks including left-right disentanglement, shape matching, and part segmentation demonstrate their effectiveness.' The body contains no such experiments. The tables (e.g., Tables 1–5, S1–S18) report object recognition, semantic segmentation, optical flow, and robustness results for event-camera methods. None of the reported metrics pertain to 3D shape analysis, chirality, or the named downstream tasks. Consequently, the abstract's empirical claims are unsupported.
- [Technical Assumptions (unaddressed)] Even if the intended chirality method were present, the abstract's premise that 2D foundation models encode handedness cues and that per-vertex aggregation preserves them is nontrivial. The manuscript does not discuss horizontal-flip augmentation (commonly used in 2D pretraining, which removes left/right bias) or the need for a consistent global reference to define 'left' vs. 'right' across different shapes. Since none of this is addressed anywhere in the submitted text, these risks are unassessed.
minor comments (3)
- [Author/Title Mismatch] The submitted header lists authors associated with the chirality paper and a project page (https://wei-kang-wang.github.io/chirality/), while the body lists Lin Zhu, Ruonan Liu, Xiao Wang, Lizhi Wang, and Hua Huang as authors of the event-camera paper. The arXiv identifiers also differ. This makes the manuscript internally inconsistent and difficult to attribute.
- [References] The reference list is entirely drawn from the event-camera literature (e.g., event-based pretraining, neuromorphic datasets). There are no references to Diff3F, 3D shape matching, point cloud features, or chirality in shape analysis, which would be expected for the claimed method.
- [Figure and Table Captions] Figures 1–9 and Tables 1–8, as well as the supplementary tables, are all captioned for event camera representation learning. None of the figures or tables illustrate chirality, vertex features, or shape symmetry. The single figure referenced in the abstract is not present in the body.
Circularity Check
No circularity can be identified: the supplied full text is a different paper (event-camera pre-training), so the chirality derivation chain is absent and no step reduces to its own inputs.
full rationale
The circularity analysis requires an actual derivation chain to inspect. The abstract of arXiv:2508.05505 describes an unsupervised chirality feature extraction pipeline built on Diff3F, with per-vertex chirality-aware features, left-right disentanglement, shape matching, and part segmentation. None of that content appears in the supplied full text. Instead, the full text is the paper 'Revealing Latent Information: A Physics-inspired Self-supervised Pre-training Framework for Noisy and Sparse Events' by Zhu et al., with its own arXiv number (2508.05507v1), ACM MM '25 copyright, event-camera abstract, sections on masked modeling, feature transition, contrastive learning, and experiments on object recognition, semantic segmentation, optical flow, and robustness. There is no equation, definition, or result in the supplied text that even mentions chirality, Diff3F, shape vertices, or left/right disentanglement. Accordingly, there is no quoted step where a 'prediction' is equivalent to a fitted input, no self-citation chain that forces a conclusion, and no ansatz smuggled in via citation. The material mismatch is a serious evidentiary problem for verifying the paper's claims, but it is not a circularity: circularity requires exhibiting a specific reduction from the paper's own derivations, and no such derivation is present. Under the hard rules, honest non-finding is required when the claimed reduction cannot be exhibited. The score is therefore 0, with no circular steps listed.
Assumptions & free parameters
assumptions (3)
- domain assumption 2D foundation models encode chirality information for 3D shapes
- domain assumption Diff3F-style per-vertex feature lifting preserves chirality
- domain assumption A consistent left/right reference exists across shapes
Cite this review
Pith. "Pith review of Symmetry Understanding of 3D Shapes via Chirality Disentanglement." pith.science (2026). https://pith.science/paper/M23XW3FE
@misc{pith2026250805505,
author = {Pith},
title = {Pith review of: Symmetry Understanding of 3D Shapes via Chirality Disentanglement},
year = {2026},
howpublished = {\url{https://pith.science/paper/M23XW3FE}},
note = {Machine review of arXiv:2508.05505}
}
read the original abstract
Chirality information (i.e. information that allows distinguishing left from right) is ubiquitous for various data modes in computer vision, including images, videos, point clouds, and meshes. While chirality has been extensively studied in the image domain, its exploration in shape analysis (such as point clouds and meshes) remains underdeveloped. Although many shape vertex descriptors have shown appealing properties (e.g. robustness to rigid-body transformations), they are often not able to disambiguate between left and right symmetric parts. Considering the ubiquity of chirality information in different shape analysis problems and the lack of chirality-aware features within current shape descriptors, developing a chirality feature extractor becomes necessary and urgent. Based on the recent Diff3F framework, we propose an unsupervised chirality feature extraction pipeline to decorate shape vertices with chirality-aware information, extracted from 2D foundation models. We evaluated the extracted chirality features through quantitative and qualitative experiments across diverse datasets. Results from downstream tasks including left-right disentanglement, shape matching, and part segmentation demonstrate their effectiveness and practical utility. Project page: https://wei-kang-wang.github.io/chirality/
Reference graph
Works this paper leans on
-
[1]
Inigo Alonso and Ana C Murillo. 2019. EV-SegNet: Semantic segmentation for event-based cameras. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops . 0–0
work page 2019
-
[2]
Arnon Amir, Brian Taba, David Berg, Timothy Melano, Jeffrey McKinstry, Carmelo Di Nolfo, Tapan Nayak, Alexander Andreopoulos, Guillaume Garreau, Marcela Mendoza, et al. 2017. A low power, fully event-based gesture recogni- tion system. In Proceedings of the IEEE conference on computer vision and pattern recognition. 7243–7252
work page 2017
-
[3]
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. 2022. Data2vec: A general framework for self-supervised learning in speech, vision and language. In International Conference on Machine Learning . PMLR, 1298–1312
work page 2022
-
[4]
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. 2021. Beit: Bert pre-training of image transformers. arXiv preprint arXiv:2106.08254 (2021)
arXiv 2021
-
[5]
Yin Bi, Aaron Chadha, Alhabib Abbas, Eirina Bourtsoulatze, and Yiannis An- dreopoulos. 2020. Graph-based spatio-temporal feature learning for neuromor- phic vision sensing. IEEE Transactions on Image Processing 29 (2020), 9084–9098
work page 2020
-
[6]
Jonathan Binas, Daniel Neil, Shih-Chii Liu, and Tobi Delbruck. 2017. DDD17: End-to-end DAVIS driving dataset. arXiv preprint arXiv:1711.01458 (2017)
arXiv 2017
-
[7]
Christian Brandli, Raphael Berner, Minhao Yang, Shih-Chii Liu, and Tobi Delbruck
-
[8]
Marco Cannici, Marco Ciccone, Andrea Romanoni, and Matteo Matteucci. 2019. Asynchronous convolutional networks for object detection in neuromorphic cameras. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. 0–0
work page 2019
Show all 73 references
-
[9]
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. 2020. Unsupervised learning of visual features by contrasting cluster assignments. Advances in neural information processing systems 33 (2020), 9912–9924
2020
-
[10]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In Interna- tional conference on machine learning . PMLR, 1597–1607
2020
-
[11]
Xinlei Chen, Saining Xie, and Kaiming He. 2021. An empirical study of training self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision . 9640–9649
2021
-
[12]
Hoonhee Cho, Hyeonseong Kim, Yujeong Chae, and Kuk-Jin Yoon. 2023. Label- free event-based object recognition via joint learning with image reconstruction from events. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 19866–19877
2023
-
[13]
Tobi Delbrück, Bernabe Linares-Barranco, Eugenio Culurciello, and Christoph Posch. 2010. Activity-driven, event-based vision sensors. In Proceedings of 2010 IEEE international symposium on circuits and systems . IEEE, 2426–2429
2010
-
[14]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition . Ieee, 248–255
2009
-
[15]
Yongjian Deng, Hao Chen, Hai Liu, and Youfu Li. 2022. A voxel graph cnn for object classification with event cameras. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1172–1181
2022
-
[16]
Yongjian Deng, Hao Chen, Bochen Xie, Hai Liu, and Youfu Li. 2023. A Dynamic Graph CNN with Cross-Representation Distillation for Event-Based Recognition. arXiv preprint arXiv:2302.04177 (2023)
2023 arXiv
-
[17]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv prepri...
2020 arXiv
-
[18]
Guillermo Gallego, Tobi Delbrück, Garrick Orchard, Chiara Bartolozzi, Brian Taba, Andrea Censi, Stefan Leutenegger, Andrew J Davison, Jörg Conradt, Kostas Daniilidis, et al. 2020. Event-based vision: A survey. IEEE transactions on pattern analysis and machine intelligence 44, ...
2020
-
[19]
Peng Gao, Teli Ma, Hongsheng Li, Ziyi Lin, Jifeng Dai, and Yu Qiao. 2022. Convmae: Masked convolution meets masked autoencoders. arXiv preprint arXiv:2205.03892 (2022)
2022 arXiv
-
[20]
Daniel Gehrig, Antonio Loquercio, Konstantinos G Derpanis, and Davide Scara- muzza. 2019. End-to-end learning of representations for asynchronous event- based data. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 5633–5643
2019
-
[21]
Mathias Gehrig, Willem Aarents, Daniel Gehrig, and Davide Scaramuzza. 2021. Dsec: A stereo event camera dataset for driving scenarios. IEEE Robotics and Automation Letters 6, 3 (2021), 4947–4954
2021
-
[22]
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. 2020. Bootstrap your own latent-a new approach to self-supervised learning. Advances in ne...
2020
-
[23]
Ryuhei Hamaguchi, Yasutaka Furukawa, Masaki Onishi, and Ken Sakurada. 2023. Hierarchical neural memory network for low latency event processing. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 22867–22876
2023
-
[24]
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick
-
[25]
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Mo- mentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 9729–9738
2020
-
[26]
Zejiang Hou, Fei Sun, Yen-Kuang Chen, Yuan Xie, and Sun-Yuan Kung. 2022. Milan: Masked image pretraining on language assisted representation. arXiv preprint arXiv:2208.06049 (2022)
2022 arXiv
-
[27]
Yuhuang Hu, Shih-Chii Liu, and Tobi Delbruck. 2021. v2e: From video frames to realistic DVS events. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 1312–1321
2021
-
[28]
Lang Huang, Shan You, Mingkai Zheng, Fei Wang, Chen Qian, and Toshihiko Ya- masaki. 2022. Green hierarchical vision transformer for masked image modeling. Advances in Neural Information Processing Systems 35 (2022), 19997–20010
2022
-
[29]
Zhicheng Huang, Xiaojie Jin, Chengze Lu, Qibin Hou, Ming-Ming Cheng, Dong- mei Fu, Xiaohui Shen, and Jiashi Feng. 2023. Contrastive masked autoencoders are stronger vision learners. IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)
2023
-
[30]
Zhenpeng Huang, Chao Li, Hao Chen, Yongjian Deng, Yifeng Geng, and Limin Wang. 2024. Data-efficient Event Camera Pre-training via Disentangled Masked Modeling. arXiv preprint arXiv:2403.00416 (2024)
2024 arXiv
-
[31]
Ziyu Jiang, Yinpeng Chen, Mengchen Liu, Dongdong Chen, Xiyang Dai, Lu Yuan, Zicheng Liu, and Zhangyang Wang. 2023. Layer grafted pre-training: Bridging contrastive learning and masked image modeling for label-efficient representations. arXiv preprint arXiv:2302.14138 (2023)
2023 arXiv
-
[32]
Junho Kim, Jaehyeok Bae, Gangin Park, Dongsu Zhang, and Young Min Kim
-
[33]
Simon Klenk, David Bonello, Lukas Koestler, Nikita Araslanov, and Daniel Cre- mers. 2024. Masked event modeling: Self-supervised pretraining for event cam- eras. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2378–2388
2024
-
[34]
Lingdong Kong, Youquan Liu, Lai Xing Ng, Benoit R Cottereau, and Wei Tsang Ooi. 2024. Openess: Event-based semantic scene understanding with open vo- cabularies. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 15686–15698
2024
-
[35]
Hongmin Li, Hanchao Liu, Xiangyang Ji, Guoqi Li, and Luping Shi. 2017. Cifar10- dvs: an event-stream dataset for object classification. Frontiers in neuroscience 11 (2017), 309
2017
-
[36]
Yijin Li, Han Zhou, Bangbang Yang, Ye Zhang, Zhaopeng Cui, Hujun Bao, and Guofeng Zhang. 2021. Graph-based asynchronous event processing for rapid object recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 934–943
2021
-
[37]
Yihan Lin, Wei Ding, Shaohua Qiang, Lei Deng, and Guoqi Li. 2021. Es-imagenet: A million event-stream classification dataset for spiking neural networks. Fron- tiers in neuroscience 15 (2021), 726582
2021
-
[38]
Haotian Liu, Guang Chen, Sanqing Qu, Yanping Zhang, Zhijun Li, Alois Knoll, and Changjun Jiang. 2023. Tma: Temporal motion aggregation for event-based optical flow. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 9685–9694
2023
-
[39]
Jonathan Long, Evan Shelhamer, and Trevor Darrell. 2015. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3431–3440
2015
-
[40]
Nico Messikommer, Daniel Gehrig, Antonio Loquercio, and Davide Scaramuzza
-
[41]
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748 (2018)
2018 arXiv
-
[42]
Garrick Orchard, Ajinkya Jayawant, Gregory K Cohen, and Nitish Thakor. 2015. Converting static image datasets to spiking neuromorphic datasets using saccades. Frontiers in neuroscience 9 (2015), 437
2015
-
[43]
Yansong Peng, Yueyi Zhang, Zhiwei Xiong, Xiaoyan Sun, and Feng Wu. 2023. Get: Group event transformer for event-based vision. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 6038–6048. �� ���� ������� ������ ����� ������� �������� ��� �� ���
2023
-
[44]
Qiang Qu, Xiaoming Chen, Yuk Ying Chung, and Yiran Shen. 2024. Evrepsl: Event-stream representation via self-supervised learning for event-based vision. IEEE Transactions on Image Processing (2024)
2024
-
[45]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[46]
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image generation. In International conference on machine learning . Pmlr, 8821–8831
2021
-
[47]
Hongwei Ren, Yue Zhou, Haotian Fu, Yulong Huang, Renjing Xu, and Bojun Cheng. 2023. Ttpoint: A tensorized point cloud network for lightweight action recognition with event cameras. In Proceedings of the 31st ACM International Conference on Multimedia. 8026–8034
2023
-
[48]
Hongwei Ren, Yue Zhou, Jiadong Zhu, Xiaopeng Lin, Haotian Fu, Yulong Huang, Yuetong Fang, Fei Ma, Hao Yu, and Bojun Cheng. 2025. Rethinking efficient and effective point-based networks for event camera classification and regression. IEEE Transactions on Pattern Analysis and Ma...
2025
-
[49]
Jason Tyler Rolfe. 2016. Discrete variational autoencoders. arXiv preprint arXiv:1609.02200 (2016)
2016 arXiv
-
[50]
Simon Schaefer, Daniel Gehrig, and Davide Scaramuzza. 2022. Aegnn: Asyn- chronous event-based graph neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 12371–12381
2022
-
[51]
Amos Sironi, Manuele Brambilla, Nicolas Bourdis, Xavier Lagorce, and Ryad Benosman. 2018. HATS: Histograms of averaged time surfaces for robust event- based object classification. InProceedings of the IEEE conference on computer vision and pattern recognition. 1731–1740
2018
-
[52]
Zhaoning Sun, Nico Messikommer, Daniel Gehrig, and Davide Scaramuzza. 2022. Ess: Learning event-based semantic segmentation from still images. In European Conference on Computer Vision . Springer, 341–357
2022
-
[53]
Chenxin Tao, Xizhou Zhu, Weijie Su, Gao Huang, Bin Li, Jie Zhou, Yu Qiao, Xiaogang Wang, and Jifeng Dai. 2023. Siamese image modeling for self-supervised vision representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2132–2141
2023
-
[54]
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Advances in neural information processing systems 35 (2022), 10078–10093
2022
-
[55]
Zhexiong Wan, Yuchao Dai, and Yuxin Mao. 2022. Learning dense and continuous optical flow from an event camera. IEEE Transactions on Image Processing 31 (2022), 7237–7251
2022
-
[56]
Longhui Wei, Lingxi Xie, Wengang Zhou, Houqiang Li, and Qi Tian. 2022. Mvp: Multimodality-guided visual pre-training. In European conference on computer vision. Springer, 337–353
2022
-
[57]
Zhenzhi Wu, Hehui Zhang, Yihan Lin, Guoqi Li, Meng Wang, and Ye Tang. 2021. Liaf-net: Leaky integrate and analog fire network for lightweight and efficient spatiotemporal information processing. IEEE Transactions on Neural Networks and Learning Systems 33, 11 (2021), 6249–6262
2021
-
[58]
Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. 2018. Uni- fied perceptual parsing for scene understanding. In Proceedings of the European conference on computer vision (ECCV) . 418–434
2018
-
[59]
Bochen Xie, Yongjian Deng, Zhanpeng Shao, Hai Liu, and Youfu Li. 2022. Vmv- gcn: Volumetric multi-view based graph cnn for event stream classification. IEEE Robotics and Automation Letters 7, 2 (2022), 1976–1983
2022
-
[60]
Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu. 2022. Simmim: A simple framework for masked image modeling. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9653–9663
2022
-
[61]
Yan Yang, Liyuan Pan, and Liu Liu. 2023. Event camera data pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 10699– 10709
2023
-
[62]
Yan Yang, Liyuan Pan, and Liu Liu. 2024. Event camera data dense pre-training. In European Conference on Computer Vision . Springer, 292–310
2024
-
[63]
Man Yao, Huanhuan Gao, Guangshe Zhao, Dingheng Wang, Yihan Lin, Zhaoxu Yang, and Guoqi Li. 2021. Temporal-wise attention spiking neural networks for event streams classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 10221–10230
2021
-
[64]
Kun Yi, Yixiao Ge, Xiaotong Li, Shusheng Yang, Dian Li, Jianping Wu, Ying Shan, and Xiaohu Qie. 2022. Masked image modeling with denoising contrast. arXiv preprint arXiv:2205.09616 (2022)
2022 arXiv
-
[65]
Zongyou Yu, Qiang Qu, Qian Zhang, Nan Zhang, and Xiaoming Chen. 2025. Llm- evrep: Learning an llm-compatible event representation using a self-supervised framework. In Companion Proceedings of the ACM on Web Conference 2025 . 2314– 2319
2025
-
[66]
Xiaoyu Yue, Shuyang Sun, Zhanghui Kuang, Meng Wei, Philip HS Torr, Wayne Zhang, and Dahua Lin. 2021. Vision transformer with progressive sampling. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 387–396
2021
-
[67]
Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. 2021. ibot: Image bert pre-training with online tokenizer.arXiv preprint arXiv:2111.07832 (2021)
2021 arXiv
-
[68]
Jiazhou Zhou, Xu Zheng, Yuanhuiyi Lyu, and Lin Wang. 2024. Eventbind: Learn- ing a unified representation to bind them all for event-based open-world under- standing. In European Conference on Computer Vision . Springer, 477–494
2024
-
[69]
Alex Zihao Zhu, Dinesh Thakur, Tolga Özaslan, Bernd Pfrommer, Vijay Kumar, and Kostas Daniilidis. 2018. The multivehicle stereo event camera dataset: An event camera dataset for 3D perception. IEEE Robotics and Automation Letters 3, 3 (2018), 2032–2039. ������������� �������� ...
2018
-
[2014]
IEEE Journal of Solid-State Circuits 49, 10 (2014), 2333–2341
A 240× 180 130 db 3 �s latency global shutter spatiotemporal vision sensor. IEEE Journal of Solid-State Circuits 49, 10 (2014), 2333–2341
2014
-
[2020]
In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VIII 16
Event-based asynchronous sparse convolutional networks. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VIII 16 . Springer, 415–431
2020
-
[2021]
In Proceedings of the IEEE/CVF international conference on computer vision
N-imagenet: Towards robust, fine-grained object recognition with event cameras. In Proceedings of the IEEE/CVF international conference on computer vision. 2146–2156
-
[2022]
In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 16000–16009
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.