Presents the first large-scale infrared off-road dataset and a flow-free temporal model achieving state-of-the-art freespace detection performance with real-time inference.
hub
Imagenet: A large-scale hierarchical image database
26 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
dataset 1polarities
use dataset 1representative citing papers
Unlearnable examples fail under pretraining-finetuning due to semantic filtering by frozen layers, but Shallow Semantic Camouflage restores effectiveness by confining perturbations to semantically valid subspaces.
Floating-point weight formats in embedded neural networks suffer near-total accuracy loss from a single electromagnetic fault injection, while 8-bit integer formats retain substantially higher accuracy on the same hardware.
MINT implements dynamic-precision CNN inference on FPGA via MSDF digit-serial arithmetic and greedy per-layer precision search, reporting up to 82% higher energy efficiency than INT8 on VGG-16 and ResNet-18 with under 2% accuracy loss.
An end-to-end spiking encoder-decoder network achieves 92.05/87.04/86.51 AP on KITTI BEV detection with a claimed 3.33x synaptic energy reduction versus an equivalent CNN.
AIR amortizes 2D Gaussian splatting into a self-supervised feed-forward network via residual stages, explicit stage control, and Predict-Optimize-Distill training.
ELSA is a near-SRAM dataflow architecture realizing elastic inference in SNNs via fine-grained spine/token pipelines, bundled AER, and mini-batch Gustavson products, delivering up to 3.4x speedup and 22.1x energy gains over SOTA accelerators on ResNet-50.
This work provides the first systematic study of transferring direct-coded spiking neural networks to event-based representations while aiming to preserve accuracy and reduce energy use.
A new chain of lightweight neural predictors with information inheritance achieves near state-of-the-art lossless compression ratios while delivering 1.2-6.3x faster encoding and 2.8-12.3x faster decoding than PAC on GPUs.
LOCALUT delivers 1.82x geometric mean speedup for quantized DNN inference on real UPMEM DRAM-PIM devices by using operation-packed LUTs with canonicalization, reordering, and slice streaming.
BicKD introduces a bilateral contrastive loss in knowledge distillation that strengthens class-wise orthogonality and intra-class consistency in predictive distributions, outperforming prior logit-based methods.
Fed-TaLoRA uses task-agnostic low-rank residual adaptation with post-aggregation calibration to enable efficient federated continual fine-tuning across sequential tasks under non-IID conditions.
AAND is a two-stage anomaly detection method that advances a pre-trained teacher via residual anomaly amplification and applies hard knowledge distillation in reverse distillation to achieve SOTA results on MVTecAD, VisA, and MVTec3D-RGB.
A learnable spatial bias term, added to a Noise2Noise-style multi-frame start inside an autoregressive deblurring loop, improves self-supervised defocus deblurring under low-light biased noise.
Replacement Learning replaces selected blocks in CNNs and ViTs with learnable parameter-fusion surrogates derived from adjacent layers to reduce full-depth backpropagation redundancy.
Proposes MODIAD framework with MIS scheduling solved via SMG algorithm and REC-LoRA adaptation for efficient multimodal online distributed industrial anomaly detection, reporting superior performance on MVTec 3D-AD and Eyecandies datasets.
BerLU constructs a C1-differentiable activation with Lipschitz constant 1 via Bernstein polynomial approximation, showing better performance and efficiency than baselines on image classification with ViTs and CNNs.
Meta-ensemble learning on diverse ICBHI data splits reaches 66.49% Score and improves generalization on two external datasets.
AnyUser translates free-form sketches on images plus optional language into executable robot actions for domestic tasks using multimodal fusion and a hierarchical policy.
GAPL anchors text prompts to second-order Gram matrix statistics to improve vision-language model adaptation across domains.
TEA is a new targeted adversarial attack that incorporates edge information from the target image to reduce query count and improve performance in low-query black-box hard-label settings.
Multi-horizon forecasting with deep learning on sky images and PV data improves prediction accuracy across multiple future time steps and architectures by jointly optimizing sequences of outputs.
Introduces circulate-firing neurons, time-step-wise learnable surrogate gradients, and balanced loss for direct SNN training, reporting competitive results on datasets and Transformers.
DACP is a proposed protocol with Streaming Data Frame as its core model that enables cross-domain scientific data discovery, in-situ computation, and streaming result return.
citing papers explorer
-
Towards All-Day Perception for Off-Road Driving: A Large-Scale Multispectral Dataset and Comprehensive Benchmark
Presents the first large-scale infrared off-road dataset and a flow-free temporal model achieving state-of-the-art freespace detection performance with real-time inference.
-
Channel-Level Semantic Perturbations: Unlearnable Examples for Diverse Training Paradigms
Unlearnable examples fail under pretraining-finetuning due to semantic filtering by frozen layers, but Shallow Semantic Camouflage restores effectiveness by confining perturbations to semantically valid subspaces.
-
The Weight of a Bit: EMFI Sensitivity Analysis of Embedded Deep Learning Models
Floating-point weight formats in embedded neural networks suffer near-total accuracy loss from a single electromagnetic fault injection, while 8-bit integer formats retain substantially higher accuracy on the same hardware.
-
MINT: Dynamic-Precision CNN Inference with MSDF Digit-Serial Arithmetic on FPGA
MINT implements dynamic-precision CNN inference on FPGA via MSDF digit-serial arithmetic and greedy per-layer precision search, reporting up to 82% higher energy efficiency than INT8 on VGG-16 and ResNet-18 with under 2% accuracy loss.
-
Neuromorphic LiDAR-based Bird's Eye View Object Detection using Energy-efficient Spiking Neural Networks
An end-to-end spiking encoder-decoder network achieves 92.05/87.04/86.51 AP on KITTI BEV detection with a claimed 3.33x synaptic energy reduction versus an equivalent CNN.
-
AIR: Amortized Image Reconstruction Framework for Self-Supervised Feed-Forward 2D Gaussian Splatting
AIR amortizes 2D Gaussian splatting into a self-supervised feed-forward network via residual stages, explicit stage control, and Predict-Optimize-Distill training.
-
ELSA: An ELastic SNN Inference Architecture for Efficient Neuromorphic Computing
ELSA is a near-SRAM dataflow architecture realizing elastic inference in SNNs via fine-grained spine/token pipelines, bundled AER, and mini-batch Gustavson products, delivering up to 3.4x speedup and 22.1x energy gains over SOTA accelerators on ResNet-50.
-
Direct-to-Event Spiking Neural Network Transfer
This work provides the first systematic study of transferring direct-coded spiking neural networks to event-based representations while aiming to preserve accuracy and reduce energy use.
-
Lossless Compression via Chained Lightweight Neural Predictors with Information Inheritance
A new chain of lightweight neural predictors with information inheritance achieves near state-of-the-art lossless compression ratios while delivering 1.2-6.3x faster encoding and 2.8-12.3x faster decoding than PAC on GPUs.
-
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
LOCALUT delivers 1.82x geometric mean speedup for quantized DNN inference on real UPMEM DRAM-PIM devices by using operation-packed LUTs with canonicalization, reordering, and slice streaming.
-
BicKD: Bilateral Contrastive Knowledge Distillation
BicKD introduces a bilateral contrastive loss in knowledge distillation that strengthens class-wise orthogonality and intra-class consistency in predictive distributions, outperforming prior logit-based methods.
-
Task-agnostic Low-rank Residual Adaptation for Efficient Federated Continual Fine-Tuning
Fed-TaLoRA uses task-agnostic low-rank residual adaptation with post-aggregation calibration to enable efficient federated continual fine-tuning across sequential tasks under non-IID conditions.
-
Advancing Pre-trained Teacher: Towards Robust Feature Discrepancy for Anomaly Detection
AAND is a two-stage anomaly detection method that advances a pre-trained teacher via residual anomaly amplification and applies hard knowledge distillation in reverse distillation to achieve SOTA results on MVTecAD, VisA, and MVTec3D-RGB.
-
Physen-Noise2Noise: Physics-Guided Self-Supervised Defocus Deblurring with Bias Correction under Low-Light Conditions
A learnable spatial bias term, added to a Noise2Noise-style multi-frame start inside an autoregressive deblurring loop, improves self-supervised defocus deblurring under low-light biased noise.
-
Replacement Learning: Training Neural Networks with Fewer Parameters
Replacement Learning replaces selected blocks in CNNs and ViTs with learnable parameter-fusion surrogates derived from adjacent layers to reduce full-depth backpropagation redundancy.
-
Parameter Efficient Multi-Class Intelligent Scheduling for Multimodal Online Distributed Industrial Anomaly Detection
Proposes MODIAD framework with MIS scheduling solved via SMG algorithm and REC-LoRA adaptation for efficient multimodal online distributed industrial anomaly detection, reporting superior performance on MVTec 3D-AD and Eyecandies datasets.
-
Universal Smoothness via Bernstein Polynomials: A Constructive Approximation Approach for Activation Functions
BerLU constructs a C1-differentiable activation with Lipschitz constant 1 via Bernstein polynomial approximation, showing better performance and efficiency than baselines on image classification with ViTs and CNNs.
-
Meta-Ensemble Learning with Diverse Data Splits for Improved Respiratory Sound Classification
Meta-ensemble learning on diverse ICBHI data splits reaches 66.49% Score and improves generalization on two external datasets.
-
AnyUser: Translating Sketched User Intent into Domestic Robots
AnyUser translates free-form sketches on images plus optional language into executable robot actions for domestic tasks using multimodal fusion and a hierarchical policy.
-
Gram-Anchored Prompt Learning for Vision-Language Models via Second-Order Statistics
GAPL anchors text prompts to second-order Gram matrix statistics to improve vision-language model adaptation across domains.
-
Accelerating Targeted Hard-Label Adversarial Attacks in Low-Query Black-Box Settings
TEA is a new targeted adversarial attack that incorporates edge information from the target image to reduce query count and improve performance in low-query black-box hard-label settings.
-
Learning Long-Term Temporal Dependencies in Photovoltaic Power Output Prediction Through Multi-Horizon Forecasting
Multi-horizon forecasting with deep learning on sky images and PV data improves prediction accuracy across multiple future time steps and architectures by jointly optimizing sequences of outputs.
-
Advancing Direct Training for Spiking Neural Networks with Circulate-Firing Neurons and Learnable Gradients
Introduces circulate-firing neurons, time-step-wise learnable surrogate gradients, and balanced loss for direct SNN training, reporting competitive results on datasets and Transformers.
-
DACP: A Scientific Data Access and Collaboration Protocol
DACP is a proposed protocol with Streaming Data Frame as its core model that enables cross-domain scientific data discovery, in-situ computation, and streaming result return.
-
Neural Network Optimization Reimagined: Decoupled Techniques for Scratch and Fine-Tuning
DualOpt decouples optimization by using real-time layer-wise weight decay for scratch training and weight rollback for fine-tuning to improve convergence, generalization, and reduce knowledge forgetting.
- EmbodiTTA: Resource-Efficient Test-Time Adaptation for Embodied Visual Systems