Pith. sign in

REVIEW 3 cited by

Learning Robust Global Representations by Penalizing Local Predictive Power

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.13549 v2 pith:HHTC2FUO submitted 2019-05-29 cs.CV

classification cs.CV
keywords predictivelocalnetworkspowerconvolutionaldomainglobalmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite their renowned predictive power on i.i.d. data, convolutional neural networks are known to rely more on high-frequency patterns that humans deem superficial than on low-frequency patterns that agree better with intuitions about what constitutes category membership. This paper proposes a method for training robust convolutional networks by penalizing the predictive power of the local representations learned by earlier layers. Intuitively, our networks are forced to discard predictive signals such as color and texture that can be gleaned from local receptive fields and to rely instead on the global structures of the image. Across a battery of synthetic and benchmark domain adaptation tasks, our method confers improved generalization out of the domain. Also, to evaluate cross-domain transfer, we introduce ImageNet-Sketch, a new dataset consisting of sketch-like images, that matches the ImageNet classification validation set in categories and scale.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Dense scaling-law fits across model sizes, datasets and tasks show that MaMMUT (contrastive plus captioning loss) outperforms standard CLIP at large compute scales, with a consistent crossover around 1e10 to 1e11 GFLOPs.

  2. Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Class-name-free accuracy of LLM descriptors is 15.5% versus 59.5% with class names; selecting attributes from target images raises attribute-only accuracy to 23.8% (45.5% uncapped).

  3. Revisiting Bayesian Model Averaging in the Era of Foundation Models

    cs.LG 2025-05 reject novelty 4.0 of 10

    The paper proposes Bayesian model averaging and an entropy-minimizing weight optimizer for ensembling foundation models, reporting accuracy gains over output averaging on image and text classification tasks.

Pith tools