REVIEW 11 cited by
Approximating CNNs with Bag-of-local-Features models works surprisingly well on ImageNet
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Deep Neural Networks (DNNs) excel on many complex perceptual tasks but it has proven notoriously difficult to understand how they reach their decisions. We here introduce a high-performance DNN architecture on ImageNet whose decisions are considerably easier to explain. Our model, a simple variant of the ResNet-50 architecture called BagNet, classifies an image based on the occurrences of small local image features without taking into account their spatial ordering. This strategy is closely related to the bag-of-feature (BoF) models popular before the onset of deep learning and reaches a surprisingly high accuracy on ImageNet (87.6% top-5 for 33 x 33 px features and Alexnet performance for 17 x 17 px features). The constraint on local features makes it straight-forward to analyse how exactly each part of the image influences the classification. Furthermore, the BagNets behave similar to state-of-the art deep neural networks such as VGG-16, ResNet-152 or DenseNet-169 in terms of feature sensitivity, error distribution and interactions between image parts. This suggests that the improvements of DNNs over previous bag-of-feature classifiers in the last few years is mostly achieved by better fine-tuning rather than by qualitatively different decision strategies.
Forward citations
Cited by 11 Pith papers
-
Cumulative Meta-Learning from Active Learning Queries for Robustness to Spurious Correlations
CAML meta-learns a progressively refined inductive bias from active-learning queries to improve robustness to spurious correlations, reporting accuracy gains on minority groups across several benchmarks.
-
On the Reliability of Cue Conflict and Beyond
Stylized cue-conflict bias scores are confounded by impure cues, imbalance, ratio metrics and restricted labels; REFINED-BIAS supplies pure balanced stimuli and full-label MRR sensitivity for reliable diagnosis.
-
AIM: Amending Inherent Interpretability via Self-Supervised Masking
AIM uses multi-stage feature guidance for self-supervised masking to improve both interpretability (EPG) and accuracy on vision benchmarks.
-
Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
Low within-subdataset diversity and large between-subdataset differences cause shortcut learning in generalist robot policies, and targeted augmentation can mitigate it.
-
Err on the Side of Texture: Texture Bias on Real Data
A new metric and texture-identification method show that ImageNet classifiers rely heavily on specific textures, but the claim that texture bias explains natural adversarial examples is largely a consequence of how te...
-
LEMUR 2: Unlocking Neural Network Diversity for AI
LEMUR 2 releases a multi-generator, multi-task neural-architecture corpus with real-device latency metadata intended as fuel for LLM-driven AutoML.
-
Same Compression Principle, Different Geometry: Rate-Distortion Signatures Dissociate Biological and Artificial Visual Systems
RD geometry derived from confusion matrices separates humans from deep vision models, but the signatures are properties of a fitted cost matrix rather than measured trade-offs.
-
SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation
A GAN framework that translates unlabeled medical images between classes and fuses ensemble, time-averaged pseudo-labels outperforms six prior GAN semi-supervised methods on MedMNIST at 5-50 labels per class.
-
Feature-Enhanced TResNet for Fine-Grained Food Image Classification
A TResNet variant with style recalibration and criss-cross-style attention reports modest Top-1 accuracy gains on two Chinese food benchmarks.
-
Predicting Visual Memory Schemas with Variational Autoencoders
Variational autoencoders generate higher-resolution dual-channel visual memory schema maps that separately predict true and false memorability, extending prior CNN approaches.
-
Object Learning and Robust 3D Reconstruction
The thesis presents FlowCapsules, RobustNeRF, and SpotLessSplats, demonstrating that unsupervised object-based learning with motion and geometric consistency improves segmentation and 3D reconstruction.
Discussion (0). Continue with ORCID to comment.