Wake Vision pipeline produces a 6M-image person detection dataset for TinyML with 2.2% label error, improving model accuracy up to 6.6% over prior VWW benchmark across architectures and subsets.
What do compressed deep neural networks forget?arXiv preprint arXiv:1911.05248
10 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 3representative citing papers
Systematic experiments reveal that activation steering trades fluency for concept control, is less effective on instruction-tuned models, and that prompting/SFT excel at injection but not removal, with textual metrics correlating to LLM judges.
Sigma-Branch converts dense networks into single-path hierarchical trees via activation-based spherical k-means clustering and soft-routing fine-tuning, cutting active parameters 58-60% with under 2pp accuracy loss on CIFAR-100, ImageNet-1K, and ModelNet40.
3-bit quantization induces new stereotypical biases in 6-21% of previously unbiased BBQ items across three LLMs, undetected by perplexity increases under 3%, with models declining in 'unknown' responses by 17.4%.
SalUn uses gradient-based weight saliency to achieve effective machine unlearning of data, classes, or concepts in image classification and generation, narrowing the gap to exact retraining.
DPIFrame introduces intra- and inter-module parallel architecture, anticipatory multi-table embedding lookup, and breadth-first GPU stream scheduling, reporting 23x embedding latency reduction and up to 5.83x overall speedup versus PyTorch on real datasets.
Activation-aware pruning preserves perplexity but amplifies bias in LLMs, with 47-59% of previously neutral items developing new stereotypical responses at 70% sparsity.
BISE extracts bias-free subnetworks from conventionally trained models via pruning, enabling debiased operation without retraining or additional data.
Token compression in ViT segmentation degrades sharply at high ratios due to information loss while structural pruning degrades smoothly; a moderate prune-then-merge pipeline improves the trade-off on ADE20K and Cityscapes under corruption.
citing papers explorer
-
Wake Vision: A Tailored Dataset and Benchmark Suite for TinyML Computer Vision Applications
Wake Vision pipeline produces a 6M-image person detection dataset for TinyML with 2.2% label error, improving model accuracy up to 6.6% over prior VWW benchmark across architectures and subsets.
-
On The Effectiveness-Fluency Trade-Off In LLM Conditioning: A Systematic Study
Systematic experiments reveal that activation steering trades fluency for concept control, is less effective on instruction-tuned models, and that prompting/SFT excel at injection but not removal, with textual metrics correlating to LLM judges.
-
Sigma-Branch: Hierarchical Single-Path Network Reconstruction for Dynamic Inference with Reduced Active Parameters
Sigma-Branch converts dense networks into single-path hierarchical trees via activation-based spherical k-means clustering and soft-routing fine-tuning, cutting active parameters 58-60% with under 2pp accuracy loss on CIFAR-100, ImageNet-1K, and ModelNet40.
-
Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels
3-bit quantization induces new stereotypical biases in 6-21% of previously unbiased BBQ items across three LLMs, undetected by perplexity increases under 3%, with models declining in 'unknown' responses by 17.4%.
-
SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation
SalUn uses gradient-based weight saliency to achieve effective machine unlearning of data, classes, or concepts in image classification and generation, narrowing the gap to exact retraining.
-
DPIFrame: A Dual-Level Parallelism Acceleration Framework for CTR Model Inference
DPIFrame introduces intra- and inter-module parallel architecture, anticipatory multi-table embedding lookup, and breadth-first GPU stream scheduling, reporting 23x embedding latency reduction and up to 5.83x overall speedup versus PyTorch on real datasets.
-
Weight Pruning Amplifies Bias: A Multi-Method Study of Compressed LLMs for Edge AI
Activation-aware pruning preserves perplexity but amplifies bias in LLMs, with 47-59% of previously neutral items developing new stereotypical responses at 70% sparsity.
-
Bias In, Bias Out? Finding Unbiased Subnetworks in Vanilla Models
BISE extracts bias-free subnetworks from conventionally trained models via pruning, enabling debiased operation without retraining or additional data.
-
When Token Compression Breaks: Structural Pruning vs. Token Reduction for Robust ViT Segmentation under High Compression
Token compression in ViT segmentation degrades sharply at high ratios due to information loss while structural pruning degrades smoothly; a moderate prune-then-merge pipeline improves the trade-off on ADE20K and Cityscapes under corruption.
- Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI