REVIEW 6 cited by
Robust and Generalizable Visual Representation Learning via Random Convolutions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Robust and Generalizable Visual Representation Learning via Random Convolutions
read the original abstract
While successful for various computer vision tasks, deep neural networks have shown to be vulnerable to texture style shifts and small perturbations to which humans are robust. In this work, we show that the robustness of neural networks can be greatly improved through the use of random convolutions as data augmentation. Random convolutions are approximately shape-preserving and may distort local textures. Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local textures. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training. When applying a network trained with our approach to unseen domains, our method consistently improves the performance on domain generalization benchmarks and is scalable to ImageNet. In particular, in the challenging scenario of generalizing to the sketch domain in PACS and to ImageNet-Sketch, our method outperforms state-of-art methods by a large margin. More interestingly, our method can benefit downstream tasks by providing a more robust pretrained visual representation.
Forward citations
Cited by 6 Pith papers
-
CAD-Free Learning of Spacecraft Pose Estimators via NeRF-Based Augmentations
NeRF augmentation trains accurate spacecraft pose estimators from 25-400 real images without CAD models or large synthetic datasets.
-
CAD-Free Learning of Spacecraft Pose Estimators via NeRF-Based Augmentations
NeRF-based image augmentation enables accurate target-specific spacecraft pose estimators to be trained from only 25-400 real images without CAD models or large synthetic datasets.
-
Frequency Adapter with SAM for Generalized Medical Image Segmentation
FSAM integrates a frequency adapter into SAM with LoRA to extract domain-invariant high-frequency features and outperforms prior domain generalization methods on fundus and prostate datasets.
-
Why Invariance is Not Enough for Biomedical Domain Generalization and How to Fix It
MaskGen improves domain generalization for biomedical image segmentation by using source intensities plus domain-stable foundation model representations with minimal added complexity.
-
One Sequence to Segment Them All: Efficient Data Augmentation for CT and MRI Cross-Domain 3D Spine Segmentation
Targeted data augmentations let single-sequence 3D spine segmentation models generalize to seven unseen CT and MRI datasets with 155% average Dice gain and almost no in-domain loss.
-
FGML-DG: Feynman-Inspired Cognitive Science Paradigm for Cross-Domain Medical Image Segmentation
FGML-DG applies Feynman-inspired principles of concept simplification, memory recall, and error-focused retraining within a meta-learning setup to enhance domain generalization for medical image segmentation.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.