Domain-specific macOS features enable an ML detector to reach 98.5% accuracy on 41k samples and 99.5% on 9k fresh samples, beating prior methods by 16-50%.
Ishaan Watts, Catherine Li, Sachin Goyal, Jacob Mitchell Springer, and Aditi Raghunathan
11 Pith papers cite this work, alongside 794 external citations. Polarity classification is still indexing.
representative citing papers
Introduces a unified benchmark for continual anomaly detection with discrete and continuous protocols plus a training-free DINOSaur method that outperforms prior CAD approaches with zero forgetting and sub-100ms edge inference.
Inf-SSM constrains the infinite-horizon evolution of SSMs via Grassmannian geometry and an efficient O(n^2) Sylvester solver to enable exemplar-free continual learning with reduced forgetting.
DSINet uses a selective spatial state unit (S3U) from Mamba and concentration-balanced distillation (CBD) to keep spatial change representations stable across incremental domains while preserving linear efficiency.
Acquisition route affects forgetting rates in multimodal models, with text-pathway knowledge forgetting faster than audio-pathway knowledge in music understanding tasks.
SynLearner lets LLMs improve synthetic data generation on later tasks in a stream by learning reusable patterns and balancing quality with diversity from feedback on earlier tasks.
Pretraining LR decay sharpens LLMs, and that sharpness—not just token count—drives catastrophic forgetting during supervised fine-tuning.
Proposes LoRA-based mixture-of-experts with autoencoder routing for continual bidirectional motion-language learning, reporting near-zero forgetting on a 5-task HumanML3D benchmark derived via semantic clustering.
A validation-gated multi-agent framework enables online adaptation of thermal-hydraulic surrogates and reduces forecast error by 19% under regime shifts on experimental loop data.
Proposes the CBDT framework as a minimum viable digital twin for CI builds to enable real-time monitoring, ML modeling, and prescriptive optimization of build duration, failures, and flakiness.
This survey defines the Federated Continual Learning problem, proposes a taxonomy for approaches, reviews applications and metrics, and identifies open challenges in lifelong privacy-preserving learning on non-stationary distributed data.
citing papers explorer
-
The Role of Domain-Specific Features in Malware Detection: A macOS Case Study
Domain-specific macOS features enable an ML detector to reach 98.5% accuracy on 41k samples and 99.5% on 9k fresh samples, beating prior methods by 16-50%.
-
Rethinking Continual Anomaly Detection on the Edge: Benchmarking Under Realistic Industrial Conditions
Introduces a unified benchmark for continual anomaly detection with discrete and continuous protocols plus a training-free DINOSaur method that outperforms prior CAD approaches with zero forgetting and sub-100ms edge inference.
-
Exemplar-Free Continual Learning for State Space Models
Inf-SSM constrains the infinite-horizon evolution of SSMs via Grassmannian geometry and an efficient O(n^2) Sylvester solver to enable exemplar-free continual learning with reduced forgetting.
-
Dual-Selective Network for Domain-Incremental Change Detection
DSINet uses a selective spatial state unit (S3U) from Mamba and concentration-balanced distillation (CBD) to keep spatial change representations stable across incremental domains while preserving linear efficiency.
-
When the Same Musical Knowledge Forgets Differently: A Clean Probe of Pathway-Dependent Forgetting
Acquisition route affects forgetting rates in multimodal models, with text-pathway knowledge forgetting faster than audio-pathway knowledge in music understanding tasks.
-
Make LLM Learn to Synthesize from Streaming Experiences through Feedback
SynLearner lets LLMs improve synthetic data generation on later tasks in a stream by learning reusable patterns and balancing quality with diversity from feedback on earlier tasks.
-
(How) Learning Rates Regulate Catastrophic Overtraining
Pretraining LR decay sharpens LLMs, and that sharpness—not just token count—drives catastrophic forgetting during supervised fine-tuning.
-
Towards Continual Motion-Language Agents: LoRA Variants for Incremental Motion Understanding and Generation
Proposes LoRA-based mixture-of-experts with autoencoder routing for continual bidirectional motion-language learning, reporting near-zero forgetting on a 5-task HumanML3D benchmark derived via semantic clustering.
-
Validation-Gated Multi-Agent Governance for Online Adaptation of Thermal-Hydraulic Surrogate Models under Operating-Regime Shift
A validation-gated multi-agent framework enables online adaptation of thermal-hydraulic surrogates and reduces forecast error by 19% under regime shifts on experimental loop data.
-
Towards Build Optimization Using Digital Twins
Proposes the CBDT framework as a minimum viable digital twin for CI builds to enable real-time monitoring, ML modeling, and prescriptive optimization of build duration, failures, and flakiness.
-
Federated continual learning: A comprehensive survey on lifelong and privacy-preserving learning over distributed and non-stationary data
This survey defines the Federated Continual Learning problem, proposes a taxonomy for approaches, reviews applications and metrics, and identifies open challenges in lifelong privacy-preserving learning on non-stationary distributed data.