Q-Margin encodes margin penalties into the reference measure of an alpha-divergence loss to produce sparse discriminative embeddings for face and speaker verification.
hub Canonical reference
Benign Overfitting in Binary Classification of Gaussian Mix- tures
Canonical reference. 80% of citing Pith papers cite this work as background.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
Sarashina2.2-TTS achieves SOTA kanji reading accuracy via data scaling and Joyo-kanji-targeted synthesis, introduces the Joyo Kanji Yomi Benchmark and Kana-CER metric, and shows stable cross-lingual performance.
Predictive Entropy Maximization performs competitive blind source separation using only local error-driven and Hebbian updates derived from a surrogate entropy objective with spectral error bounds.
X-VC achieves zero-shot streaming voice conversion via one-step codec-space conversion with dual-conditioning acoustic converter and role-assignment training on generated paired data.
ORBGRAND-AI achieves the same or lower block error rate in ISI channels without interleaving compared to CA-SCL decoding with an interleaver at equal energy per information bit.
New conditions for support vector proliferation (SVP) in RKHS for bounded orthonormal systems and sub-Gaussian features, yielding generalization bounds for kernel SVMs beyond prior restrictive assumptions.
Secure-CHG introduces a cascaded defense with statistical filtering early and CHG-Shapley valuation later to mitigate late-stage failure against backdoor attacks in federated learning, reporting 2.3x and 2.0x lower attack success rates than Krum and Trimmed Mean on CIFAR-10, MedMNIST, and NEU-SDDB.
Retrieval from out-of-domain foundation models enables personalization of a lightweight transformer for stress detection, yielding +3.92% accuracy and +4.76% F1 gains on WESAD without user labels.
Derives smoothness-based PAC-Bayes derandomization bounds for deterministic predictors using Rademacher complexity of the Jensen gap class, yielding Jacobian/Hessian flatness terms and a practical regularizer tested on CIFAR-10.
Presents Streaming-Train-248K dataset, Streaming Harness system, and Streaming-Eval benchmark to enable VLMs for proactive, memory-equipped streaming video understanding.
AdaCodec introduces a predictive visual code that cuts visual token use in video MLLMs by sending full frames only on high predictive cost and otherwise encoding inter-frame changes as P-tokens, yielding better benchmark scores at lower budgets.
SwanVoice is a zero-shot TTS system for 1-4 speakers that reports higher richness and hierarchy scores than open-source baselines on monologue and dialogue tasks via mixed training and DiffusionNFT post-training.
Proposes a psychovisual-inspired deep learning method that encodes images in learned frequency sub-bands for interpretable semantic structures and reduced depth dependence.
Shared-score quaternion self-attention reduces score multiplications by 75% and softmax operations from four to one while proving equivalence to component-wise attention under quaternion linear projections.
Reinforcement learning with graph neural networks finds minimally rigid graphs that match known planar realization optima and set new records for spherical realization counts.
An encoding probe reconstructs transformer representations from acoustic, phonetic, syntactic, lexical and speaker features, showing independent syntactic/lexical contributions and training-dependent speaker effects.
A co-design framework using approximate matrix decomposition and genetic algorithms delivers 33% average latency reduction in TinyML CNN FPGA accelerators with 1.3% average accuracy loss versus standard systolic arrays.
Minimizing the sum of ℓ∞ norms enables separation of antisparse bounded sources via PCA followed by Givens rotations optimization, with claimed superior performance over prior methods in simulations.
A framework using covariance-based spectral signatures and TreeSHAP attributions on AASIST3 branches identifies four operational archetypes and a flawed specialization mode that explains high error rates on specific spoofing attacks.
Latent SDE generative model for anomaly detection in sparse irregular multivariate time series outperforms baselines on six benchmarks and stays robust under severe sparsity.
HP-VSR-ResFiLM adds a single residual FiLM modulation block conditioned on head pose to a CNN visual encoder, yielding WER of 25.0% on LRS2 and 33.2% on LRS3 under standard training conditions.
The RER framework decomposes chord generation into retrieval, editing, and reranking stages to outperform end-to-end models in balancing stylistic diversity with music-theoretic feasibility.
Hybrid RL-PID controllers track angle of attack better and show greater robustness than PID alone within a defined operational envelope for re-entry attitude control.
Deep GLR combines graph Laplacian regularization with three lightweight CNN modules in a proximal optimization framework to reach 30.70 dB PSNR on LoDoPaB-CT using 5.8x fewer parameters and 30x less data per dB gain than typical deep methods.
citing papers explorer
-
Co-Design of CNN Accelerators for TinyML using Approximate Matrix Decomposition
A co-design framework using approximate matrix decomposition and genetic algorithms delivers 33% average latency reduction in TinyML CNN FPGA accelerators with 1.3% average accuracy loss versus standard systolic arrays.