Pith. sign in

REVIEW 11 cited by

Approximating CNNs with Bag-of-local-Features models works surprisingly well on ImageNet

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.00760 v1 pith:BUMLJQ7W submitted 2019-03-20 cs.CV cs.LGstat.ML

classification cs.CVcs.LGstat.ML
keywords featuresimagedeepimagenetarchitecturebag-of-featuredecisionsdnns
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Deep Neural Networks (DNNs) excel on many complex perceptual tasks but it has proven notoriously difficult to understand how they reach their decisions. We here introduce a high-performance DNN architecture on ImageNet whose decisions are considerably easier to explain. Our model, a simple variant of the ResNet-50 architecture called BagNet, classifies an image based on the occurrences of small local image features without taking into account their spatial ordering. This strategy is closely related to the bag-of-feature (BoF) models popular before the onset of deep learning and reaches a surprisingly high accuracy on ImageNet (87.6% top-5 for 33 x 33 px features and Alexnet performance for 17 x 17 px features). The constraint on local features makes it straight-forward to analyse how exactly each part of the image influences the classification. Furthermore, the BagNets behave similar to state-of-the art deep neural networks such as VGG-16, ResNet-152 or DenseNet-169 in terms of feature sensitivity, error distribution and interactions between image parts. This suggests that the improvements of DNNs over previous bag-of-feature classifiers in the last few years is mostly achieved by better fine-tuning rather than by qualitatively different decision strategies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cumulative Meta-Learning from Active Learning Queries for Robustness to Spurious Correlations

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    CAML meta-learns a progressively refined inductive bias from active-learning queries to improve robustness to spurious correlations, reporting accuracy gains on minority groups across several benchmarks.

  2. On the Reliability of Cue Conflict and Beyond

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Stylized cue-conflict bias scores are confounded by impure cues, imbalance, ratio metrics and restricted labels; REFINED-BIAS supplies pure balanced stimuli and full-label MRR sensitivity for reliable diagnosis.

  3. AIM: Amending Inherent Interpretability via Self-Supervised Masking

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    AIM uses multi-stage feature guidance for self-supervised masking to improve both interpretability (EPG) and accuracy on vision benchmarks.

  4. Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation

    cs.RO 2025-08 conditional novelty 6.0 of 10

    Low within-subdataset diversity and large between-subdataset differences cause shortcut learning in generalist robot policies, and targeted augmentation can mitigate it.

  5. Err on the Side of Texture: Texture Bias on Real Data

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new metric and texture-identification method show that ImageNet classifiers rely heavily on specific textures, but the claim that texture bias explains natural adversarial examples is largely a consequence of how te...

  6. LEMUR 2: Unlocking Neural Network Diversity for AI

    cs.LG 2026-07 conditional novelty 5.5 of 10

    LEMUR 2 releases a multi-generator, multi-task neural-architecture corpus with real-device latency metadata intended as fuel for LLM-driven AutoML.

  7. Same Compression Principle, Different Geometry: Rate-Distortion Signatures Dissociate Biological and Artificial Visual Systems

    cs.LG 2026-03 reject novelty 5.0 of 10

    RD geometry derived from confusion matrices separates humans from deep vision models, but the signatures are properties of a fitted cost matrix rather than measured trade-offs.

  8. SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    A GAN framework that translates unlabeled medical images between classes and fuses ensemble, time-averaged pseudo-labels outperforms six prior GAN semi-supervised methods on MedMNIST at 5-50 labels per class.

  9. Feature-Enhanced TResNet for Fine-Grained Food Image Classification

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A TResNet variant with style recalibration and criss-cross-style attention reports modest Top-1 accuracy gains on two Chinese food benchmarks.

  10. Predicting Visual Memory Schemas with Variational Autoencoders

    cs.CV 2019-07 unverdicted novelty 4.0 of 10

    Variational autoencoders generate higher-resolution dual-channel visual memory schema maps that separately predict true and false memorability, extending prior CNN approaches.

  11. Object Learning and Robust 3D Reconstruction

    cs.CV 2025-04 accept novelty 2.0 of 10

    The thesis presents FlowCapsules, RobustNeRF, and SpotLessSplats, demonstrating that unsupervised object-based learning with motion and geometric consistency improves segmentation and 3D reconstruction.

Pith tools