EQ-VMamba adds rotation-equivariant cross-scan and group Mamba blocks to enforce end-to-end rotation equivariance, yielding better rotation robustness, competitive accuracy, and roughly 50% fewer parameters than non-equivariant baselines across classification, segmentation, and super-resolution.
super hub Mixed citations
Very Deep Convolutional Networks for Large-Scale Image Recognition
Mixed citation behavior. Most common role is background (58%).
abstract
In this work we investigate the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting. Our main contribution is a thorough evaluation of networks of increasing depth using an architecture with very small (3x3) convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers. These findings were the basis of our ImageNet Challenge 2014 submission, where our team secured the first and the second places in the localisation and classification tracks respectively. We also show that our representations generalise well to other datasets, where they achieve state-of-the-art results. We have made our two best-performing ConvNet models publicly available to facilitate further research on the use of deep visual representations in computer vision.
hub tools
citation-role summary
citation-polarity summary
claims ledger
- abstract In this work we investigate the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting. Our main contribution is a thorough evaluation of networks of increasing depth using an architecture with very small (3x3) convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers. These findings were the basis of our ImageNet Challenge 2014 submission, where our team secured the first and the second places in the localisation and classification tracks respective
authors
co-cited works
representative citing papers
Real NVP uses affine coupling layers to create invertible transformations that support exact density estimation, sampling, and latent inference without approximations.
A u-shaped fully-convolutional encoder-decoder with skip connections trained with elastic-deformation augmentation produces accurate biomedical image segmentations from very small training sets.
The paper derives a posteriori error estimates for neural network depth adaptation by formulating training as an optimal control problem and using dual weighted residuals to insert layers where error is highest.
A near-infrared POV-based physical attack achieves high success rates fooling traffic sign classifiers across 12 models, distances to 20m, and varying conditions.
Fine-tunes EG3D using a human-preference reward on NeRF density to improve face geometry, achieving 74.4% user preference in pairwise tests with FID rising from 4.09 to 6.66.
A variational autoencoder learns quantum embeddings compressing ImageNet into 13 qubits and achieving 98.5% accuracy on MNIST 3-vs-5 classification with a quantum circuit, close to classical baselines and far above naive amplitude embeddings.
DIPBox is the first multi-scale testing framework for detecting adversarial dataset regeneration via four similarity metrics, backed by learning-theoretic analysis of utility-divergence trade-offs.
ExSpike is a full-event neuromorphic FPGA architecture using dataflow optimizations, a spike-driven attention core, and adjacent-position event compression to exploit irregular sparsity in SNNs, delivering up to 10x higher PE-normalized energy efficiency than prior SOTA on classification and segment
A diffusion-trained maximum entropy model uses 512 learned statistics to synthesize visual textures at quality matching or exceeding prior models that rely on ~177k statistics.
Dolph2Vec is the first species-specific self-supervised model for dolphin vocalizations, trained on longitudinal recordings from five dolphins, that outperforms general baselines on signature whistle classification and detection while producing embeddings aligned with known whistle categories.
A large-scale empirical study across tokenizers and diffusion backbones identifies Velocity Irreducible Variance (VIV) as one of the most stable predictors of latent diffusion generation quality.
Chameleon proposes the first large-scale cross-domain compositing dataset and a disentangled encoder plus gated diffusion transformer that outperforms prior in-domain and cross-domain methods on plausibility and fidelity.
SGDM is shown to be algorithmically stable on smooth convex problems, yielding optimal excess population risk bounds for both Polyak and Nesterov momentum.
PROBE improves AIGI detector generalization to unseen generators by using the detector as a critic to steer manifold-level modifications that produce challenging training samples.
MAPS provides 2618 validated 3D meshes and a controllable rendering pipeline to attribute vision model recognition failures to specific scene parameters, finding camera distance and elevation as the dominant failure factors across 20 tested models.
Mirror flow reaches max-margin solutions in homogeneous neural networks where the mirror map choice controls whether learned features are sparse or dense while convergence can be exponentially slow.
Next-acceleration-scale autoregressive prediction in discrete latent space with on-policy privileged information distillation yields improved MRI reconstructions from sparse measurements on the fastMRI benchmark.
A new differentiable reconstruction method uses symmetrized hyperspherical harmonics on quaternions plus two- and three-point descriptors to generate 3D microstructures from 2D data, demonstrated on aluminum alloy with L-BFGS-B optimization.
Introduces synthetic ground-truth dataset for CAM evaluation, proposes ARCC composite metric, and RefineCAM method that aggregates layers for higher-resolution maps outperforming baselines.
KamonBench is a grammar-based dataset of 20,000 synthetic Japanese crests with multi-format annotations that enables direct evaluation of factor recovery beyond caption accuracy in vision-language models.
SubPopMark embeds verifiable subpopulation biases into distilled datasets via CVM and USTM optimization stages, allowing provenance inference through comparison of model output signatures against a reference behavior bank.
Human face perception aligns with neural networks trained on inverse-generative and naturalistic discriminative tasks, as these best predict human dissimilarity judgments on controversial and random face pairs.
CoDAAR aligns modality-specific codebooks at the index level using Discrete Temporal Alignment and Cascading Semantic Alignment to achieve cross-modal generalization while preserving unique structures, reporting state-of-the-art results on event classification, localization, video segmentation, and跨
citing papers explorer
-
Rotation Equivariant Mamba for Vision Tasks
EQ-VMamba adds rotation-equivariant cross-scan and group Mamba blocks to enforce end-to-end rotation equivariance, yielding better rotation robustness, competitive accuracy, and roughly 50% fewer parameters than non-equivariant baselines across classification, segmentation, and super-resolution.
-
Density estimation using Real NVP
Real NVP uses affine coupling layers to create invertible transformations that support exact density estimation, sampling, and latent inference without approximations.
-
U-Net: Convolutional Networks for Biomedical Image Segmentation
A u-shaped fully-convolutional encoder-decoder with skip connections trained with elastic-deformation augmentation produces accurate biomedical image segmentations from very small training sets.
-
An optimal control approach for neural network architecture adaptation with a posteriori error estimation
The paper derives a posteriori error estimates for neural network depth adaptation by formulating training as an optimal control problem and using dual weighted residuals to insert layers where error is highest.
-
The Spectrum Strikes Back: Infrared POV Attacks on Traffic Sign Classification
A near-infrared POV-based physical attack achieves high success rates fooling traffic sign classifiers across 12 models, distances to 20m, and varying conditions.
-
Sculpting NeRF Geometry: Human-Preference Fine-Tuning of a 3D-Aware Face GAN
Fine-tunes EG3D using a human-preference reward on NeRF density to improve face geometry, achieving 74.4% user preference in pairwise tests with FID rising from 4.09 to 6.66.
-
Tailor Made Embeddings for Quantum Machine Learning
A variational autoencoder learns quantum embeddings compressing ImageNet into 13 qubits and achieving 98.5% accuracy on MNIST 3-vs-5 classification with a quantum circuit, close to classical baselines and far above naive amplitude embeddings.
-
DIPBox: A Multi-scale Testing Framework for Tracking Dataset Regeneration
DIPBox is the first multi-scale testing framework for detecting adversarial dataset regeneration via four similarity metrics, backed by learning-theoretic analysis of utility-divergence trade-offs.
-
ExSpike: A General Full-Event Neuromorphic Architecture for Exploiting Irregular Sparsity with Event Compression
ExSpike is a full-event neuromorphic FPGA architecture using dataflow optimizations, a spike-driven attention core, and adjacent-position event compression to exploit irregular sparsity in SNNs, delivering up to 10x higher PE-normalized energy efficiency than prior SOTA on classification and segment
-
Learning a Maximum Entropy Model for Visual Textures using Diffusion
A diffusion-trained maximum entropy model uses 512 learned statistics to synthesize visual textures at quality matching or exceeding prior models that rely on ~177k statistics.
-
Dolph2Vec: Self-Supervised Representations of Dolphin Vocalizations
Dolph2Vec is the first species-specific self-supervised model for dolphin vocalizations, trained on longitudinal recordings from five dolphins, that outperforms general baselines on signature whistle classification and detection while producing embeddings aligned with known whistle categories.
-
Diffusing in the Right Space: A Systematic Study of Latent Diffusability
A large-scale empirical study across tokenizers and diffusion backbones identifies Velocity Irreducible Variance (VIV) as one of the most stable predictors of latent diffusion generation quality.
-
Chameleon: Style-Content Disentangled Framework for Cross-Domain Object Compositing
Chameleon proposes the first large-scale cross-domain compositing dataset and a disentangled encoder plus gated diffusion transformer that outperforms prior in-domain and cross-domain methods on plausibility and fidelity.
-
Stochastic Gradient Descent with Momentum is Algorithmically Stable
SGDM is shown to be algorithmically stable on smooth convex problems, yielding optimal excess population risk bounds for both Polyak and Nesterov momentum.
-
Where Detectors Fail: Probing Generative Space for Generalizable AI-Generated Image Detection
PROBE improves AIGI detector generalization to unseen generators by using the detector as a critic to steer manifold-level modifications that produce challenging training samples.
-
MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
MAPS provides 2618 validated 3D meshes and a controllable rendering pipeline to attribute vision model recognition failures to specific scene parameters, finding camera distance and elevation as the dominant failure factors across 20 tested models.
-
Implicit Bias of Mirror Flow in Homogeneous Neural Networks: Sparse and Dense Feature Learning
Mirror flow reaches max-margin solutions in homogeneous neural networks where the mirror map choice controls whether learned features are sparse or dense while convergence can be exponentially slow.
-
Next-Acceleration-Scale Prediction for Autoregressive MRI Reconstruction
Next-acceleration-scale autoregressive prediction in discrete latent space with on-policy privileged information distillation yields improved MRI reconstructions from sparse measurements on the fastMRI benchmark.
-
Generative reconstruction of 2D and 3D polycrystalline microstructures using symmetrized hyperspherical harmonics
A new differentiable reconstruction method uses symmetrized hyperspherical harmonics on quaternions plus two- and three-point descriptors to generate 3D microstructures from 2D data, demonstrated on aluminum alloy with L-BFGS-B optimization.
-
How to Evaluate and Refine your CAM
Introduces synthetic ground-truth dataset for CAM evaluation, proposes ARCC composite metric, and RefineCAM method that aggregates layers for higher-resolution maps outperforming baselines.
-
KamonBench: A Grammar-Based Dataset for Evaluating Compositional Factor Recovery in Vision-Language Models
KamonBench is a grammar-based dataset of 20,000 synthetic Japanese crests with multi-format annotations that enables direct evaluation of factor recovery beyond caption accuracy in vision-language models.
-
From Compression to Accountability: Harmless Copyright Protection for Dataset Distillation
SubPopMark embeds verifiable subpopulation biases into distilled datasets via CVM and USTM optimization stages, allowing provenance inference through comparison of model output signatures against a reference behavior bank.
-
Human face perception reflects inverse-generative and naturalistic discriminative objectives
Human face perception aligns with neural networks trained on inverse-generative and naturalistic discriminative tasks, as these best predict human dissimilarity judgments on controversial and random face pairs.
-
Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations
CoDAAR aligns modality-specific codebooks at the index level using Discrete Temporal Alignment and Cascading Semantic Alignment to achieve cross-modal generalization while preserving unique structures, reporting state-of-the-art results on event classification, localization, video segmentation, and跨
-
TCP-SSM: Efficient Vision State Space Models with Token-Conditioned Poles
TCP-SSM conditions stable poles on visual tokens to explicitly control memory decay and oscillation in SSMs, cutting computation up to 44% while matching or exceeding accuracy on classification, segmentation, and detection.
-
Empirical Evidence for Simply Connected Decision Regions in Image Classifiers
Empirical tests with quad-mesh filling indicate that decision regions in modern image classifiers are simply connected.
-
Retain-Neutral Surrogates for Min-Max Unlearning
ROSU derives a closed-form retain-neutral perturbation for min-max unlearning that bounds retain damage via curvature and improves performance when gradients are aligned.
-
DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models
DMGD achieves better performance than fine-tuned SOTA methods in dataset distillation on ImageNet subsets by using semantic matching through conditional likelihood optimization and OT-based distribution matching in a training-free diffusion setup.
-
Hierarchical Spatio-Channel Clustering for Efficient Model Compression in Medical Image Analysis
A spatio-channel clustering framework for CNN compression reduces FLOPs by 81% and raises brain tumor MRI classification accuracy from 87.76% to 89.80% compared with global SVD and Tucker baselines.
-
KAConvNet: Kolmogorov-Arnold Convolutional Networks for Vision Recognition
KAConvNet introduces a Kolmogorov-Arnold Convolutional Layer to build networks competitive with ViTs and CNNs while offering stronger theoretical interpretability.
-
Different Strokes for Different Folks: Writer Identification for Historical Arabic Manuscripts
CNN models with attention reach 99.05% top-1 accuracy on line-level splits and 78.61% on page-disjoint splits for writer identification after expanding the labeled portion of the Muharaf historical Arabic manuscript dataset.
-
MESA: A Training-Free Multi-Exemplar Deep Framework for Restoring Ancient Inscription Textures
MESA restores ancient inscription textures via multi-exemplar style transfer from VGG19 features with per-layer exemplar selection and OCR-derived weights, without any model training.
-
Channel-Level Semantic Perturbations: Unlearnable Examples for Diverse Training Paradigms
Unlearnable examples fail under pretraining-finetuning due to semantic filtering by frozen layers, but Shallow Semantic Camouflage restores effectiveness by confining perturbations to semantically valid subspaces.
-
Physically-Induced Atmospheric Adversarial Perturbations: Enhancing Transferability and Robustness in Remote Sensing Image Classification
FogFool creates fog-based adversarial perturbations using Perlin noise optimization to achieve high black-box transferability (83.74% TASR) and robustness to defenses in remote sensing classification.
-
VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale
VidTAG achieves fine-grained global video-to-GPS geolocalization via temporal frame alignment and denoising sequence refinement, reporting 20% gains at 1 km over GeoCLIP and 25% on CityGuessr68k.
-
Beyond Corner Patches: Semantics-Aware Backdoor Attack in Federated Learning
SABLE shows that semantics-aware natural triggers enable effective backdoor attacks in federated learning against multiple aggregation rules while preserving benign accuracy.
-
Do Machines Fail Like Humans? A Human-Centred Out-of-Distribution Spectrum for Mapping Error Alignment
A human-centered OOD spectrum based on perceptual difficulty shows vision-language models align best with human errors across regimes, with CNNs stronger on near-OOD and ViTs on far-OOD.
-
Retinex Meets Language: A Physics-Semantics-Guided Underwater Image Enhancement Network
PSG-UIENet fuses Retinex physics with CLIP-derived text semantics and a new multimodal dataset to enhance underwater images, claiming better results than fifteen prior methods.
-
FedBCD:Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated Learning
FedBCGD reduces communication in federated learning by a factor of 1/N through block-wise parameter updates with accelerated convergence guarantees.
-
Two-stage Convolutional Neural Network for pseudo six-dimensional phase space reconstruction
A two-stage CNN reconstructs pseudo 6D phase space from 16 x-y images taken at varying rotation angles in the KEK-ATF injector.
-
Descriptor: Parasitoid Wasps and Associated Hymenoptera Dataset (DAPWH)
Releases the DAPWH dataset of 3556 wasp images including 1739 COCO-annotated examples to enable AI models for identifying Ichneumonoidea and associated families.
-
The Weight of a Bit: EMFI Sensitivity Analysis of Embedded Deep Learning Models
Floating-point weight formats in embedded neural networks suffer near-total accuracy loss from a single electromagnetic fault injection, while 8-bit integer formats retain substantially higher accuracy on the same hardware.
-
A Case for Hypergraphs to Model and Map SNNs on Neuromorphic Hardware
Hypergraph modeling of SNNs improves neuron-to-core mapping on neuromorphic hardware by exploiting hyperedge overlap and locality for better partitioning and placement than graph-based methods.
-
Thinking Like Van Gogh: Structure-Aware Style Transfer via Flow-Guided 3D Gaussian Splatting
Flow-guided advection in 3D Gaussian Splatting transfers 2D artistic motion into 3D geometry to produce structure-aware stylization.
-
LooseRoPE: Content-aware Attention Manipulation for Semantic Harmonization
LooseRoPE modulates RoPE in diffusion attention maps to continuously trade off between preserving a pasted object's identity and harmonizing it with its new surroundings.
-
Fusion2Print: Deep Flash-Non-Flash Fusion for Contactless Fingerprint Matching
Fusion2Print fuses flash-non-flash contactless fingerprints via attention-based networks and U-Net enhancement to reach AUC 0.999 and EER 1.12% with cross-domain compatibility.
-
Agentic AI in Remote Sensing: Foundations, Taxonomy, and Emerging Systems
The paper delivers the first comprehensive review and unified taxonomy of agentic AI in remote sensing, covering single-agent copilots, multi-agent systems, planning mechanisms, benchmarks, and a roadmap while noting limitations in grounding and safety.
-
It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models
Noise optimization during sampling recovers diversity in mode-collapsed diffusion models while preserving output fidelity.
-
GenCellAgent: Generalizable, Training-Free Cellular Image Segmentation via Large Language Model Agents
GenCellAgent deploys a planner-executor-evaluator LLM agent loop to automatically select, adapt, and refine segmentation tools for diverse cellular microscopy images, matching or exceeding specialist performance on 4,718 images across seven benchmarks while handling out-of-distribution and novel-ves
-
A fast machine learning tool to predict the composition of astronomical ices from infrared absorption spectra
Neural-network model trained on lab ice spectra predicts fractional composition of H2O, CO, CO2, CH3OH, NH3, and CH4 from 2.5-10 micron IR absorption with typical 3% error and was validated on two JWST background-star spectra.