Pith. sign in

REVIEW 2 major objections 4 minor 73 references

Foundation Models and Transformers for Anomaly Detection: A Survey

T0 review · 2 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This survey claims that Transformers and foundation models have transformed visual anomaly detection, overcoming CNN limits and enabling zero-shot detection.

desk verdict Useful survey scope and taxonomy undermined by two placeholder arXiv references in the bias discussion; the citation base needs verification before the paper is trustworthy. read the letter →

arxiv 2507.15905 v1 pith:FIEKIPU6 submitted 2025-07-21 cs.LG cs.AI

classification cs.LGcs.AI
keywords anomalydetectionvisualTransformersfoundationmodelszero-shotself-supervisedlearningattentionmechanismsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey tries to establish that attention-based architectures and foundation models have transformed visual anomaly detection (VAD), moving the field from per-class models trained on normal examples toward pre-trained, promptable systems that detect anomalies with little or no task-specific data. It argues that the global receptive field of Transformers addresses long-range and logical anomalies that convolutional networks struggle with, and that large-scale pre-training plus vision-language alignment enables zero- and few-shot detection. To make this case it organizes published methods into reconstruction-based, feature-based, and zero-/few-shot families, including video and point-cloud work. A sympathetic reader would use this as a structured map of the field and a statement of why the field's center of gravity has shifted.

What carries the argument

The machinery that carries the argument is the attention mechanism, running in self-attention, cross-attention, and masked-attention forms, together with large-scale pre-training. Concretely, the paper relies on the softmax attention operation $\mathrm{softmax}(QK^T/\sqrt{d_k})V$, Vision Transformer patch encoders, masked-autoencoder reconstruction, contrastive vision-language alignment as in CLIP, prompt tuning and learnable prompts, and memory banks of pre-trained features. These components supply the global receptive field, the representation quality, and the zero-shot transferability that the survey says convolutional approaches lack.

What would settle it

Resolving the two arXiv identifiers '2303.12345' and '2306.12345' in the reference list would settle whether the survey's citation base is genuine; if they do not correspond to published papers, then readers cannot rely on the accuracy of the method summaries that depend on those citations.

Watch

Extended reading notes

Core claim

The paper's central claim is that Transformers and foundation models have transformed visual anomaly detection. On its account, the attention mechanism gives every token a global receptive field, letting detectors model long-range and contextual dependencies that the local receptive fields of CNNs miss. Foundation models trained on web-scale data, especially CLIP and SAM, relax the need for category-specific training: anomalies can be recognized by aligning image features with textual descriptions of normality, and segmentation masks can be produced zero-shot. The survey's distinct contribution is a taxonomy organizing these methods into reconstruction-based, feature-based, and zero-/few-shot approaches, with attention-based video and point-cloud methods included.

Load-bearing premise

The survey's central claim stands on the assumption that every cited method is a real, published work that the authors actually read and summarized faithfully, since a review's map of the field is only as reliable as its references.

Editorial extensions

If this is right

  • If the survey is right, VAD research should focus less on one-model-per-class training and more on pre-trained, prompt-adaptable systems that serve many classes at once.
  • Reconstruction-based detectors can avoid the identity-mapping trap by using query embeddings, masked reconstruction, or pre-trained feature targets rather than only deeper autoencoders.
  • Feature-based methods inheriting self-supervised ViT representations become the data-efficient default, with memory banks providing test-time reference without retraining.
  • Zero-shot anomaly detection via CLIP-style alignment will keep improving with better prompts, value-value attention, and segmentation backbones like SAM.
  • The taxonomy implies that evaluation should cover image, video, and point-cloud data, not just single-image benchmarks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the same taxonomy could be applied to other dense prediction tasks that use foundation models, such as open-set segmentation or novelty detection in point clouds, not only to visual anomaly detection.
  • Editorial inference: if foundation-model methods keep improving, evaluation practice should likely shift from per-class AUROC toward prompt robustness, calibration under distribution shift, and computational cost at deployment.
  • Editorial inference: the survey's emphasis on data scarcity suggests that research on few-shot prompt selection and memory-bank pruning would follow directly from its framework and would be testable with existing benchmarks.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This paper is a survey of visual anomaly detection (VAD) methods built on Transformer architectures and foundation models. It proposes a taxonomy with three main families—reconstruction-based, feature-based, and zero-/few-shot detection—and reviews representative methods in each, with detailed subsections on CLIP/SAM-based approaches, other vision-language models, and non-foundation-model few-shot methods. It also provides background on Transformer attention mechanisms and on the anomaly/out-of-distribution/novelty-detection terminology. The survey does not contain original experiments, and its stated contribution is to organize and synthesize the literature on the paradigm shift from CNNs to Transformers and foundation models in VAD.

Significance. The intended contribution is timely: existing surveys focus on CNNs, GANs, diffusion models, or OOD detection, while this one targets Transformer and foundation-model VAD and explicitly includes video and point-cloud methods. If the literature coverage were accurate, the taxonomy in Figure 3 and the organization of Sections 4–6 would be a useful entry point for researchers. The paper gives credit to a wide range of recent works (e.g., WinCLIP, AnomalyCLIP, PromptAD, LAVAD, SAM-LAD) and makes sensible high-level observations about identity mapping, data scarcity, prompt engineering, and interpretability. However, the survey's value rests entirely on the accuracy of its citations and method summaries; there is no experimental validation, code, or machine-checked content. The bibliography contains entries that cannot be verified, which is load-bearing for the paper's central claim.

major comments (2)
  1. [Section 6.4 and References] The reference list contains two entries with the canonical placeholder arXiv identifiers 2303.12345 and 2306.12345: 'Hao et al. Zhang, Challenges of foundation models in industrial visual inspection, arXiv:2303.12345, 2023' and 'Yujia Zhang, Tianwei Li, and Xingyu Liu, Fairness and robustness in anomaly detection: A survey, arXiv:2306.12345, 2023g'. These are not real papers. The first is cited as 'Zhang (2023)' to support the claim that foundation-model biases cause high false-positive rates in medical and industrial inspection, and the second is cited as 'Zhang et al. (2023g)' to support bias-aware evaluation. Because the survey has no original results, these citations are the only support for the affected claims; invalid citations leave the bias discussion unsupported and undermine confidence in the reliability of the other method summaries. The malformed author string 'Hao et al. Zhang' is also consistent with a template-generated entry rather than a checked source. This issue is load-bearing for the survey's central contribution and cannot be dismissed as a typo.
  2. [Bibliography (general)] Beyond the two placeholder IDs, the reference list contains several entries with malformed author strings, such as 'Mahmudul et al. Hasan', 'Mehrdad et al. Ravanbakhsh', 'Zexin et al. Wu', and 'Linchao et al. Yu'. These may be real publications with formatting errors, but the pattern means the reader cannot easily distinguish checked entries from unchecked ones. Since the paper's entire evidence base is the cited literature, the authors would need to provide an audited reference list, not just fix the two invalid IDs, before the taxonomy could be trusted.
minor comments (4)
  1. [Section 1.1.1, Eq. (2)] Equation (2) defines masked self-attention as softmax(QK^T / sqrt(dq) ∘ M) and omits the multiplication by V that appears in Eq. (1); as written, the equation is incomplete.
  2. [References] The bibliography is formatted inconsistently: many entries use the pattern 'Given-name et al. Surname' (e.g., 'Hao et al. Zhang'), which should be normalized to the standard 'Surname et al.' convention.
  3. [Section 1.1] Vaswani et al. is cited as (2023) in the text while the surrounding sentence says the model was introduced in 2017; the citation should refer to the original NeurIPS 2017 paper or the reprint should be explicitly identified.
  4. [Introduction and Section 5.3] The Introduction claims coverage of point-cloud methods, but the only point-cloud-specific method appears to be Multi-3D-Memory in Section 5.3, with no dedicated discussion of point-cloud-specific challenges such as sparsity or irregular structure.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey is an exposition of external literature with no derivation chain, no self-citations, and no fitted inputs; the placeholder arXiv identifiers are a citation-integrity concern, not a circularity concern.

full rationale

This paper is a survey, not a derivation or experimental study. Its central claim that Transformers and foundation models have transformed visual anomaly detection is supported by citing a large body of external prior work, and the paper contains no equations whose outputs are defined in terms of their inputs, no parameters fitted to data and then renamed as predictions, and no methodological results imported from the authors' own prior publications. The authors do not cite themselves in any load-bearing way; inspection of the reference list shows no entries authored by Ben Ammar, Mendoza, Belkhir, Manzanera, or Franchi. The taxonomy in Figure 3 and Table 1 is a literature-organizing scheme, and categorizing methods into reconstruction-based, feature-based, and zero-/few-shot families is an expository choice rather than a conclusion derived from the cited works by construction. I do note a serious citation-integrity problem outside the scope of circularity: the reference list contains two placeholder arXiv identifiers, '2303.12345' for 'Hao et al. Zhang, Challenges of foundation models in industrial visual inspection' and '2306.12345' for 'Yujia Zhang, Tianwei Li, and Xingyu Liu, Fairness and robustness in anomaly detection: A survey', and these are used in Section 6.4 to support claims about inherited biases in foundation models. If these references are not real, the bias discussion loses its stated evidence base, but this does not make the survey circular because the survey's claims are not equivalent to its own inputs; they are assertions about external literature whose accuracy remains an editorial verification matter, not a self-referential loop. The survey is therefore best scored 0 on the circularity scale, with the placeholder citations flagged as a separate integrity issue.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

This is a survey paper. It introduces no free parameters, no axioms beyond the standard background knowledge of the field, and no new entities. The only assumptions are that the cited methods exist and are correctly described, which is exactly what the fabricated references call into question.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Foundation Models and Transformers for Anomaly Detection: A Survey." pith.science (2026). https://pith.science/paper/FIEKIPU6

@misc{pith2026250715905,
  author       = {Pith},
  title        = {Pith review of: Foundation Models and Transformers for Anomaly Detection: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FIEKIPU6}},
  note         = {Machine review of arXiv:2507.15905}
}
read the original abstract

In line with the development of deep learning, this survey examines the transformative role of Transformers and foundation models in advancing visual anomaly detection (VAD). We explore how these architectures, with their global receptive fields and adaptability, address challenges such as long-range dependency modeling, contextual modeling and data scarcity. The survey categorizes VAD methods into reconstruction-based, feature-based and zero/few-shot approaches, highlighting the paradigm shift brought about by foundation models. By integrating attention mechanisms and leveraging large-scale pre-training, Transformers and foundation models enable more robust, interpretable, and scalable anomaly detection solutions. This work provides a comprehensive review of state-of-the-art techniques, their strengths, limitations, and emerging trends in leveraging these architectures for VAD.

Figures

Figures reproduced from arXiv: 2507.15905 by the authors.

Figure 1
Figure 1. Vision Transformer ViT Architecture (left). Mean size of attended area (receptive field) [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Illustration of the cross-attention operation. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Taxonomy of anomaly detection methods [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Overview of UNet-like reconstruction- and prediction-based AD methods. Transform [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: The basic flow of auto-encoder reconstruction and prediction based AD methods. The [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]
Figure 6
Figure 6. Figure 6: The basic setups for multi-class and single-class anomaly detection. Single-class setup [PITH_FULL_IMAGE:figures/full_fig_p015_6.png]
Figure 7
Figure 7. Figure 7: The basic flow of distillation-based methods. For training, the student is guided by [PITH_FULL_IMAGE:figures/full_fig_p017_7.png]
Figure 8
Figure 8. Figure 8: The basic flow of distribution-map, one-class classification, and memory-bank meth [PITH_FULL_IMAGE:figures/full_fig_p019_8.png]
Figure 9
Figure 9. Figure 9: The basic flow of foundation models-based zero- and few-shot methods. The image and [PITH_FULL_IMAGE:figures/full_fig_p022_9.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

73 extracted references · 37 canonical work pages

  1. [1]

    Rafiqul Islam

    Mohiuddin Ahmed, Abdun Naser Mahmood, and Md. Rafiqul Islam. A survey of anomaly de- tection techniques in financial domain. Future Generation Computer Systems , 55:278–288, 2016a. ISSN 0167-739X. doi: https://doi.org/10.1016/j.future.2015.01.001. URL https: //www.sciencedirect.com/science/article/pii/S0167739X15000023. Mohiuddin Ahmed, Abdun Naser Mahmoo...

  2. [3]

    doi: https://doi.org/10.1016/j

    ISSN 2214-7853. doi: https://doi.org/10.1016/j. matpr.2022.01.171. URL https://www.sciencedirect.com/science/article/ pii/S2214785322001997. Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Luˇ ci´ c, and Cordelia Schmid. Vivit: A video vision transformer,

  3. [5]

    Openflamingo: An open- source framework for training large autoregressive vision-language models

    Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, et al. Openflamingo: An open- source framework for training large autoregressive vision-language models. arXiv preprint arXiv:2308.01390,

  4. [9]

    On the oppor- tunities and risks of foundation models

    Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the oppor- tunities and risks of foundation models. arXiv preprint arXiv:2108.07258,

  5. [10]

    Behzad Bozorgtabar, Dwarikanath Mahapatra, and Jean-Philippe Thiran

    doi: 10.1609/aaai.v37i12.26720. Behzad Bozorgtabar, Dwarikanath Mahapatra, and Jean-Philippe Thiran. Anomaly detection and localization using attention-guided synthetic anomaly and test-time adaptation. In 33rd British Machine Vision Conference 2022, BMVC 2022, London, UK, November 21-24, 2022 . BMVA Press,

  6. [11]

    doi: https://doi.org/10.1016/j

    ISSN 0952-1976. doi: https://doi.org/10.1016/j. engappai.2023.106677. Zhi Cai, Yingjie Gao, Yaoyan Zheng, Nan Zhou, and Di Huang. Crowd-sam: Sam as a smart an- notator for object detection in crowded scenes. arXiv preprint arXiv:2407.11464,

  7. [13]

    Deep learning for anomaly detection: A survey

    Raghavendra Chalapathy and Sanjay Chawla. Deep learning for anomaly detection: A survey. arXiv preprint arXiv:1901.03407,

  8. [14]

    doi: 10.1016/j.neunet.2021.12.008

    ISSN 0893-6080. doi: 10.1016/j.neunet.2021.12.008. Pi-Wei Chen, Jerry Chun-Wei Lin, Jia Ji, Feng-Hao Yeh, Zih-Ching Chen, and Chao-Chun Chen. Human-free automated prompting for vision-language anomaly detection: Prompt optimiza- tion with meta-guiding prompt scheme, 2024a. Sachin Mehta Chen and Mohammad Rastegari. Mobilevit: Light-weight, general-purpose,...

Show all 73 references
  1. [15]

    Tevad: Im- proved video anomaly detection with captions

    Weiling Chen, Keng Teck Ma, Zi Jian Yew, Minhoe Hur, and David Aik-Aun Khoo. Tevad: Im- proved video anomaly detection with captions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5549–5559, 2023a. Xuhai Chen, Yue Han, and Jiangning ...

  2. [16]

    Matan Jacob Cohen and Shai Avidan

    doi: 10.3390/electronics11152306. Matan Jacob Cohen and Shai Avidan. Transformaly-two (feature spaces) are better than one. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 4060–4069,

  3. [17]

    doi: 10.1142/s0129065722500307

    ISSN 1793-6462. doi: 10.1142/s0129065722500307. Harm De Vries et al. Does clip solve everything? bias and robustness in vision-language models. In Findings of EMNLP,

  4. [18]

    doi: 10.5220/0011669400003417

    ISBN 978-989-758-634-7. doi: 10.5220/0011669400003417. Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszko- reit, and Neil Houlsby. An image is...

  5. [19]

    doi: https://doi.org/10.1016/j.optlastec.2023.110296

    ISSN 0030-3992. doi: https://doi.org/10.1016/j.optlastec.2023.110296. André Luiz Buarque Vieira e Silva, Francisco Simões, Danny Kowerko, Tobias Schlosser, Felipe Battisti, and Veronica Teichrieb. Attention modules improve image-level anomaly detection for industrial inspectio...

  6. [20]

    Enhancing few-shot video anomaly detection with key-frame selection and relational cross transformers

    Ahmed Fakhry and Jong Taek Lee. Enhancing few-shot video anomaly detection with key-frame selection and relational cross transformers. In 2024 IEEE International Conference on Ad- vanced Video and Signal Based Surveillance (AVSS), pp. 1–8. IEEE,

  7. [21]

    Yan Fu, Bao Yang, and Ou Ye

    doi: 10.1145/3464423. Yan Fu, Bao Yang, and Ou Ye. Spatiotemporal masked autoencoder with multi-memory and skip connections for video anomaly detection. Electronics, 13(2):353,

  8. [22]

    Filo: Zero-shot anomaly detection by fine-grained description and high-quality localization

    Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen, Hao Li, Ming Tang, and Jinqiao Wang. Filo: Zero-shot anomaly detection by fine-grained description and high-quality localization. In Proceedings of the 32nd ACM International Conference on Multimedia, pp. 2041–2049,

  9. [23]

    doi: https://doi.org/10.1016/j.eswa.2021.116429

    ISSN 0957-4174. doi: https://doi.org/10.1016/j.eswa.2021.116429. Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network,

  10. [24]

    Jongheon Jeong, Yang Zou, Taewan Kim, Dongqing Zhang, Avinash Ravichandran, and Onkar Dabeer

    doi: 10.1109/ACCESS.2023.3234745. Jongheon Jeong, Yang Zou, Taewan Kim, Dongqing Zhang, Avinash Ravichandran, and Onkar Dabeer. Winclip: Zero-/few-shot anomaly classification and segmentation,

  11. [25]

    Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V

    32604/cmc.2023.035246. Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V . Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision,

  12. [26]

    Visual prompt tuning

    Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. Visual prompt tuning. In Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal Hassner (eds.), Computer Vision – ECCV 2022 , pp. 709–727, Cham...

  13. [27]

    Mmad: The first-ever comprehensive benchmark for multimodal large language models in industrial anomaly detection

    Xi Jiang, Jian Li, Hanqiu Deng, Yong Liu, Bin-Bin Gao, Yifeng Zhou, Jialin Li, Chengjie Wang, and Feng Zheng. Mmad: The first-ever comprehensive benchmark for multimodal large language models in industrial anomaly detection. arXiv preprint arXiv:2410.09453,

  14. [28]

    org/abs/2212.05136

    URL https://arxiv. org/abs/2212.05136. Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie. Prompting visual-language mod- els for efficient video understanding,

  15. [29]

    doi: https://doi.org/10.1016/j.knosys.2023.111186

    ISSN 0950-7051. doi: https://doi.org/10.1016/j.knosys.2023.111186. Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. Transformers in vision: A survey. ACM Computing Surveys, 54(10s):1–41,

  16. [30]

    doi: 10.1145/3505244

    ISSN 1557-7341. doi: 10.1145/3505244. Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fa- had Shahbaz Khan. Maple: Multi-modal prompt learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 19113–19122,

  17. [31]

    Convolutional networks and applica- tions in vision

    Yann LeCun, Koray Kavukcuoglu, and Clement Farabet. Convolutional networks and applica- tions in vision. In Proceedings of 2010 IEEE International Symposium on Circuits and Systems, pp. 253–256,

  18. [33]

    Selformaly: Towards task-agnostic unified anomaly detection

    Yujin Lee, Harin Lim, and Hyunsoo Yoon. Selformaly: Towards task-agnostic unified anomaly detection. arXiv preprint arXiv:2307.12540,

  19. [35]

    Sagan: Skip- attention gan for anomaly detection

    Guoliang Liu, Shiyong Lan, Ting Zhang, Weikang Huang, and Wenwu Wang. Sagan: Skip- attention gan for anomaly detection. In 2021 IEEE international conference on image process- ing (ICIP), pp. 2468–2472. IEEE, 2021a. Haodong Liu, Shouqian Sun, Yujing Ren, Chao Xu, and Jie Zhou....

  20. [36]

    Deep industrial image anomaly detection: A survey

    Jiaqi Liu, Guoyang Xie, Jinbao Wang, Shangnian Li, Chengjie Wang, Feng Zheng, and Yaochu Jin. Deep industrial image anomaly detection: A survey. Machine Intelligence Research, 21(1): 104–135, 2024a. ISSN 2731-5398. doi: 10.1007/s11633-023-1459-z. URL http://dx.doi. org/10.1007...

  21. [37]

    P- tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks

    Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. P- tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXiv preprint arXiv:2110.07602, 2021b. Yang Liu, Jing Liu, Kun Yang, Bobo Ju, Siao Liu, Y...

  22. [38]

    doi: https://doi.org/10.1016/j.engappai.2023.107810

    ISSN 0952-1976. doi: https://doi.org/10.1016/j.engappai.2023.107810. Hui Lv, Chen Chen, Zhen Cui, Chunyan Xu, Yong Li, and Jian Yang. Learning normal dynamics in videos with meta prototype network. In Proceedings of the IEEE/CVF conference on computer vision and pattern recogn...

  23. [39]

    doi: 10.1109/tpami.2023.3322604

    ISSN 1939-3539. doi: 10.1109/tpami.2023.3322604. Julien Mairal, Francis Bach, Jean Ponce, and Guillermo Sapiro. Online dictionary learning for sparse coding. In Proceedings of the 26th annual international conference on machine learning, pp. 689–696,

  24. [40]

    Anomaly detection in video using predictive convo- lutional long short-term memory networks

    Jefferson Ryan Medel and Andreas Savakis. Anomaly detection in video using predictive convo- lutional long short-term memory networks. arXiv preprint arXiv:1612.00390,

  25. [42]

    Vt-adl: A vision transformer network for image anomaly detection and localization

    Pankaj Mishra, Riccardo Verk, Daniele Fornasier, Claudio Piciarelli, and Gian Luca Foresti. Vt-adl: A vision transformer network for image anomaly detection and localization. In 2021 IEEE 30th International Symposium on Industrial Electronics (ISIE) . IEEE,

  26. [43]

    2021.9576231

    doi: 10.1109/isie45552. 2021.9576231. 35 Published as a journal paper at Information Fusion Atsuyuki Miyai, Jingkang Yang, Jingyang Zhang, Yifei Ming, Yueqian Lin, Qing Yu, Go Irie, Shafiq Joty, Yixuan Li, Hai Li, Ziwei Liu, Toshihiko Yamasaki, and Kiyoharu Aizawa. Generalized...

  27. [44]

    Zachary Novack, Julian McAuley, Zachary C

    URL https://arxiv.org/abs/2407.21794. Zachary Novack, Julian McAuley, Zachary C. Lipton, and Saurabh Garg. Chils: Zero-shot image classification with hierarchical label sets,

  28. [45]

    Accessed: 2023-03-20. OpenAI. GPT-4: Openai’ s multimodal large language model. https://openai.com/ research/gpt-4,

  29. [46]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, and et al

    Accessed: 2023-03-20. OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, and et al. Gpt-4 technical report,

  30. [47]

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al

    doi: 10.1145/3569219.3569352. Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv...

  31. [48]

    Yatian Pang, Wenxiao Wang, Francis E

    doi: 10.1088/1742-6596/2638/1/012005. Yatian Pang, Wenxiao Wang, Francis E. H. Tay, Wei Liu, Yonghong Tian, and Li Yuan. Masked autoencoders for point cloud self-supervised learning,

  32. [50]

    Very deep convolutional networks for large-scale image recognition

    Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556,

  33. [51]

    Large language models for forecasting and anomaly detection: A systematic literature review

    Jing Su, Chufeng Jiang, Xin Jin, Yuxin Qiao, Tingsong Xiao, Hongda Ma, Rong Wei, Zhi Jing, Ji- ajun Xu, and Junhong Lin. Large language models for forecasting and anomaly detection: A systematic literature review. arXiv preprint arXiv:2402.10350,

  34. [52]

    Masato Tamura

    doi: 10.1080/08839514.2022.2094885. Masato Tamura. Random word data augmentation with clip for zero-shot anomaly detection. arXiv preprint arXiv:2308.11119,

  35. [53]

    Training data-efficient image transformers & distillation through attention, 2021a

    Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention, 2021a. Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou. Go...

  36. [54]

    doi: 10.1109/TCSVT .2024. 3376399. Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake Hechtman, and Jonathon Shlens. Scaling local self-attention for parameter efficient visual backbones,

  37. [55]

    Cogvlm: Visual expert for pretrained language models

    Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, et al. Cogvlm: Visual expert for pretrained language models. arXiv preprint arXiv:2311.03079, 2023a. Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, D...

  38. [56]

    Anodfdnet: A deep feature difference network for anomaly detection, 2022b

    Zhixue Wang, Yu Zhang, Lin Luo, and Nan Wang. Anodfdnet: A deep feature difference network for anomaly detection, 2022b. Dong-Lai Wei, Chen-Geng Liu, Yang Liu, Jing Liu, Xiao-Guang Zhu, and Xin-Hua Zeng. Look, listen and pay more attention: Fusing multi-modal information for v...

  39. [58]

    Zexin et al

    URL https://arxiv.org/abs/2308.11681. Zexin et al. Wu. Denoising diffusion-augmented hybrid video anomaly detection via reconstruct- ing noised frames. In AAAI,

  40. [59]

    Limits to visual representational correspondence be- tween convolutional neural networks and the human brain

    Yaoda Xu and Maryam Vaziri-Pashkam. Limits to visual representational correspondence be- tween convolutional neural networks and the human brain. Nature communications, 12(1): 2065,

  41. [60]

    Attention-based mis- aligned spatiotemporal auto-encoder for video anomaly detection

    Haiyan Yang, Shuning Liu, Mingxuan Wu, Hongbin Chen, and Delu Zeng. Attention-based mis- aligned spatiotemporal auto-encoder for video anomaly detection. Signal, Image and Video Processing, pp. 1–13, 2024a. Hui-Yue Yang, Hui Chen, Lihao Liu, Zijia Lin, Kai Chen, Liejun Wang, J...

  42. [61]

    The dawn of lmms: Preliminary explorations with gpt-4v (ision)

    Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Li- juan Wang. The dawn of lmms: Preliminary explorations with gpt-4v (ision). arXiv preprint arXiv:2309.17421, 9(1):1, 2023a. 40 Published as a journal paper at Information Fusion Zhiwei Ya...

  43. [62]

    Visual anomaly detection via dual-attention trans- former and discriminative flow, 2023a

    Haiming Yao, Wei Luo, and Wenyong Yu. Visual anomaly detection via dual-attention trans- former and discriminative flow, 2023a. Haiming Yao, Wenyong Yu, Wei Luo, Zhenfeng Qiang, Donghao Luo, and Xiaotian Zhang. Learning global-local correspondence with semantic bottleneck for ...

  44. [63]

    Linchao et al. Yu. Stnmamba: Mamba-based spatial-temporal normality learning for video anomaly detection. arXiv preprint arXiv:2403.01234,

  45. [64]

    Luca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang, and Elisa Ricci

    doi: 10.1109/ACCESS.2021.3109102. Luca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang, and Elisa Ricci. Harness- ing large language models for training-free video anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognitio...

  46. [65]

    Visualizing and understanding convolutional networks

    Matthew D Zeiler and Rob Fergus. Visualizing and understanding convolutional networks. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part I 13, pp. 818–833. Springer,

  47. [66]

    Faster segment anything: Towards lightweight sam for mobile applica- tions

    Chaoning Zhang, Dongshen Han, Yu Qiao, Jung Uk Kim, Sung-Ho Bae, Seungkyu Lee, and Choong Seon Hong. Faster segment anything: Towards lightweight sam for mobile applica- tions. arXiv preprint arXiv:2306.14289, 2023a. Chaoning Zhang, Chenshuang Zhang, Sheng Zheng, Yu Qiao, Chen...

  48. [67]

    Exploring plain vit reconstruction for multi-class unsupervised anomaly detection, 2023c

    Jiangning Zhang, Xuhai Chen, Yabiao Wang, Chengjie Wang, Yong Liu, Xiangtai Li, Ming-Hsuan Yang, and Dacheng Tao. Exploring plain vit reconstruction for multi-class unsupervised anomaly detection, 2023c. Jiangning Zhang, Xuhai Chen, Zhucun Xue, Yabiao Wang, Chengjie Wang, and ...

  49. [68]

    doi: https://doi.org/10.1016/j.vrih.2022.07.006

    ISSN 2096-5796. doi: https://doi.org/10.1016/j.vrih.2022.07.006. Qianqian Zhang, Hongyang Wei, Jiaying Chen, Xusheng Du, and Jiong Yu. Video anomaly detec- tion based on attention mechanism. Symmetry, 15(2):528, 2023e. Shuo Zhang and Jing Liu. Feature-constrained and attention...

  50. [69]

    Improved anomaly detection based on loss prediction

    Wenkang Zhang and Fengqian Pang. Improved anomaly detection based on loss prediction. In 2022 4th International Conference on Communications, Information System and Computer En- gineering (CISCE), pp. 15–19,

  51. [70]

    Xuan Zhang, Shiyu Li, Xi Li, Ping Huang, Jiulong Shan, and Ting Chen

    doi: 10.1109/CISCE55963.2022.9850974. Xuan Zhang, Shiyu Li, Xi Li, Ping Huang, Jiulong Shan, and Ting Chen. Destseg: Segmentation guided denoising student-teacher for anomaly detection. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pp....

  52. [71]

    Yu, and Lichao Sun

    Ce Zhou, Qian Li, Chen Li, Jun Yu, Yixin Liu, Guangjing Wang, Kai Zhang, Cheng Ji, Qiben Yan, Lifang He, Hao Peng, Jianxin Li, Jia Wu, Ziwei Liu, Pengtao Xie, Caiming Xiong, Jian Pei, Philip S. Yu, and Lichao Sun. A comprehensive survey on pretrained foundation models: A histo...

  53. [72]

    doi: https://doi.org/10.1016/j.measurement.2024.114216

    ISSN 0263-2241. doi: https://doi.org/10.1016/j.measurement.2024.114216. Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592,

  54. [73]

    doi: https: //doi.org/10.1016/j.eswa.2022.118269

    ISSN 0957-4174. doi: https: //doi.org/10.1016/j.eswa.2022.118269. 43

  55. [2009]

    Paul Bergmann, Michael Fauser, David Sattlegger, and Carsten Steger

    doi: 10.1561/2200000006. Paul Bergmann, Michael Fauser, David Sattlegger, and Carsten Steger. Uninformed students: Student-teacher anomaly detection with discriminative latent embeddings. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE,

  56. [2010]

    Yann LeCun, Yoshua Bengio, and Geoffrey Hinton

    doi: 10.1109/ISCAS.2010.5537907. Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. nature, 521(7553):436–444,

  57. [2016]

    Mobilevitv2: Learning general-purpose vision repre- sentations for mobile devices

    Sachin Mehta and Mohammad Rastegari. Mobilevitv2: Learning general-purpose vision repre- sentations for mobile devices. In arXiv preprint arXiv:2206.02680,

  58. [2018]

    Cvt: Introducing convolutions to vision transformers

    Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 22–31, 2021a. Jhih-Ciang Wu, Ding-Jie Chen, Chiou-Shann F...

  59. [2019]

    Sam-lad: Segment anything model meets zero-shot logic anomaly detection

    Yun Peng, Xiao Lin, Nachuan Ma, Jiayuan Du, Chuangwei Liu, Chengju Liu, and Qijun Chen. Sam-lad: Segment anything model meets zero-shot logic anomaly detection. arXiv preprint arXiv:2406.00625,

  60. [2020]

    BigScience Workshop

    doi: 10.1109/ cvpr42600.2020.00424. BigScience Workshop. BLOOM (revision 4ab0472),

  61. [2021]

    A-vae: Attention based variational autoencoder for traffic video anomaly detection

    Nazia Aslam and Maheshkumar H Kolekar. A-vae: Attention based variational autoencoder for traffic video anomaly detection. In 2023 IEEE 8th International Conference for Convergence in Technology (I2CT), pp. 1–7. IEEE,

  62. [2022]

    Mind the pad–cnns can develop blind spots

    Bilal Alsallakh, Narine Kokhlikyan, Vivek Miglani, Jun Yuan, and Orion Reblitz-Richardson. Mind the pad–cnns can develop blind spots. arXiv preprint arXiv:2010.02178,

  63. [2023]

    Addressing ethical risks of foundation models

    Jack Bandy and Nicholas Vincent. Addressing ethical risks of foundation models. arXiv preprint arXiv:2108.08420,

  64. [2024]

    2nd place win- ning solution for the cvpr2023 visual anomaly and novelty detection challenge: Multimodal prompting for data-centric anomaly detection, 2023a

    Yunkang Cao, Xiaohao Xu, Chen Sun, Yuqi Cheng, Liang Gao, and Weiming Shen. 2nd place win- ning solution for the cvpr2023 visual anomaly and novelty detection challenge: Multimodal prompting for data-centric anomaly detection, 2023a. Yunkang Cao, Xiaohao Xu, Chen Sun, Xiaonan ...

  65. [2025]

    Mobile-former: Bridging mobilenet and transformer

    Xiangxiang Li, Wenhai Wang, Xiaoliang Zhang, Mengchen Xu, Yu Qiao, Hongyang Li, and Chun- hua Shen. Mobile-former: Bridging mobilenet and transformer. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 12328–12337, 2022b. Xiaofan Li...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.