Q-Margin encodes margin penalties into the reference measure of an alpha-divergence loss to produce sparse discriminative embeddings for face and speaker verification.
hub Canonical reference
Benign Overfitting in Binary Classification of Gaussian Mix- tures
Canonical reference. 80% of citing Pith papers cite this work as background.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
Sarashina2.2-TTS achieves SOTA kanji reading accuracy via data scaling and Joyo-kanji-targeted synthesis, introduces the Joyo Kanji Yomi Benchmark and Kana-CER metric, and shows stable cross-lingual performance.
Predictive Entropy Maximization performs competitive blind source separation using only local error-driven and Hebbian updates derived from a surrogate entropy objective with spectral error bounds.
X-VC achieves zero-shot streaming voice conversion via one-step codec-space conversion with dual-conditioning acoustic converter and role-assignment training on generated paired data.
ORBGRAND-AI achieves the same or lower block error rate in ISI channels without interleaving compared to CA-SCL decoding with an interleaver at equal energy per information bit.
New conditions for support vector proliferation (SVP) in RKHS for bounded orthonormal systems and sub-Gaussian features, yielding generalization bounds for kernel SVMs beyond prior restrictive assumptions.
Secure-CHG introduces a cascaded defense with statistical filtering early and CHG-Shapley valuation later to mitigate late-stage failure against backdoor attacks in federated learning, reporting 2.3x and 2.0x lower attack success rates than Krum and Trimmed Mean on CIFAR-10, MedMNIST, and NEU-SDDB.
Retrieval from out-of-domain foundation models enables personalization of a lightweight transformer for stress detection, yielding +3.92% accuracy and +4.76% F1 gains on WESAD without user labels.
Derives smoothness-based PAC-Bayes derandomization bounds for deterministic predictors using Rademacher complexity of the Jensen gap class, yielding Jacobian/Hessian flatness terms and a practical regularizer tested on CIFAR-10.
Presents Streaming-Train-248K dataset, Streaming Harness system, and Streaming-Eval benchmark to enable VLMs for proactive, memory-equipped streaming video understanding.
AdaCodec introduces a predictive visual code that cuts visual token use in video MLLMs by sending full frames only on high predictive cost and otherwise encoding inter-frame changes as P-tokens, yielding better benchmark scores at lower budgets.
SwanVoice is a zero-shot TTS system for 1-4 speakers that reports higher richness and hierarchy scores than open-source baselines on monologue and dialogue tasks via mixed training and DiffusionNFT post-training.
Proposes a psychovisual-inspired deep learning method that encodes images in learned frequency sub-bands for interpretable semantic structures and reduced depth dependence.
Shared-score quaternion self-attention reduces score multiplications by 75% and softmax operations from four to one while proving equivalence to component-wise attention under quaternion linear projections.
Reinforcement learning with graph neural networks finds minimally rigid graphs that match known planar realization optima and set new records for spherical realization counts.
APC embeds compact Ed25519 signatures into audio phase data with error correction to achieve 97.5-98.3% cryptographic verification under eight attack types at mean PESQ 3.02.
An encoding probe reconstructs transformer representations from acoustic, phonetic, syntactic, lexical and speaker features, showing independent syntactic/lexical contributions and training-dependent speaker effects.
A co-design framework using approximate matrix decomposition and genetic algorithms delivers 33% average latency reduction in TinyML CNN FPGA accelerators with 1.3% average accuracy loss versus standard systolic arrays.
Minimizing the sum of ℓ∞ norms enables separation of antisparse bounded sources via PCA followed by Givens rotations optimization, with claimed superior performance over prior methods in simulations.
A framework using covariance-based spectral signatures and TreeSHAP attributions on AASIST3 branches identifies four operational archetypes and a flawed specialization mode that explains high error rates on specific spoofing attacks.
Latent SDE generative model for anomaly detection in sparse irregular multivariate time series outperforms baselines on six benchmarks and stays robust under severe sparsity.
HP-VSR-ResFiLM adds a single residual FiLM modulation block conditioned on head pose to a CNN visual encoder, yielding WER of 25.0% on LRS2 and 33.2% on LRS3 under standard training conditions.
The RER framework decomposes chord generation into retrieval, editing, and reranking stages to outperform end-to-end models in balancing stylistic diversity with music-theoretic feasibility.
Hybrid RL-PID controllers track angle of attack better and show greater robustness than PID alone within a defined operational envelope for re-entry attitude control.
citing papers explorer
-
Sparsity-Inducing Divergence Losses for Biometric Verification
Q-Margin encodes margin penalties into the reference measure of an alpha-divergence loss to produce sparse discriminative embeddings for face and speaker verification.
-
Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis
Sarashina2.2-TTS achieves SOTA kanji reading accuracy via data scaling and Joyo-kanji-targeted synthesis, introduces the Joyo Kanji Yomi Benchmark and Kana-CER metric, and shows stable cross-lingual performance.
-
Normative Networks for Source Separation via Local Plasticity and Dendritic Computation
Predictive Entropy Maximization performs competitive blind source separation using only local error-driven and Hebbian updates derived from a surrogate entropy objective with spectral error bounds.
-
X-VC: Zero-shot Streaming Voice Conversion in Codec Space
X-VC achieves zero-shot streaming voice conversion via one-step codec-space conversion with dual-conditioning acoustic converter and role-assignment training on generated paired data.
-
Decoding in the presence of ISI without interleaving -- ORBGRAND-AI
ORBGRAND-AI achieves the same or lower block error rate in ISI channels without interleaving compared to CA-SCL decoding with an interleaver at equal energy per information bit.
-
New Equivalences Between Interpolation and SVMs: Kernels and Structured Features
New conditions for support vector proliferation (SVP) in RKHS for bounded orthonormal systems and sub-Gaussian features, yielding generalization bounds for kernel SVMs beyond prior restrictive assumptions.
-
Secure-CHG: A Comprehensive Framework for Robust and Fair Federated Learning via Hybrid Defense and Contribution-Aware Trust
Secure-CHG introduces a cascaded defense with statistical filtering early and CHG-Shapley valuation later to mitigate late-stage failure against backdoor attacks in federated learning, reporting 2.3x and 2.0x lower attack success rates than Krum and Trimmed Mean on CIFAR-10, MedMNIST, and NEU-SDDB.
-
Retrieval-Augmented Personalization with Foundation Models for Wearable Stress Detection
Retrieval from out-of-domain foundation models enables personalization of a lightweight transformer for stress detection, yielding +3.92% accuracy and +4.76% F1 gains on WESAD without user labels.
-
Smoothness-Based Derandomization of PAC-Bayes Bounds
Derives smoothness-based PAC-Bayes derandomization bounds for deterministic predictors using Rademacher complexity of the Jensen gap class, yielding Jacobian/Hessian flatness terms and a practical regularizer tested on CIFAR-10.
-
Harnessing Streaming Video in the Wild
Presents Streaming-Train-248K dataset, Streaming Harness system, and Streaming-Eval benchmark to enable VLMs for proactive, memory-equipped streaming video understanding.
-
AdaCodec: A Predictive Visual Code for Video MLLMs
AdaCodec introduces a predictive visual code that cuts visual token use in video MLLMs by sending full frames only on high predictive cost and otherwise encoding inter-frame changes as P-tokens, yielding better benchmark scores at lower budgets.
-
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue
SwanVoice is a zero-shot TTS system for 1-4 speakers that reports higher richness and hierarchy scores than open-source baselines on monologue and dialogue tasks via mixed training and DiffusionNFT post-training.
-
Deep Psychovisual Image Representations
Proposes a psychovisual-inspired deep learning method that encodes images in learned frequency sub-bands for interpretable semantic structures and reduced depth dependence.
-
Quaternion Self-Attention with Shared Scores
Shared-score quaternion self-attention reduces score multiplications by 75% and softmax operations from four to one while proving equivalence to component-wise attention under quaternion linear projections.
-
Learning Minimally Rigid Graphs with High Realization Counts
Reinforcement learning with graph neural networks finds minimally rigid graphs that match known planar realization optima and set new records for spherical realization counts.
-
Asymmetric Phase Coding Audio Watermarking
APC embeds compact Ed25519 signatures into audio phase data with error correction to achieve 97.5-98.3% cryptographic verification under eight attack types at mean PESQ 3.02.
-
Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe
An encoding probe reconstructs transformer representations from acoustic, phonetic, syntactic, lexical and speaker features, showing independent syntactic/lexical contributions and training-dependent speaker effects.
-
Co-Design of CNN Accelerators for TinyML using Approximate Matrix Decomposition
A co-design framework using approximate matrix decomposition and genetic algorithms delivers 33% average latency reduction in TinyML CNN FPGA accelerators with 1.3% average accuracy loss versus standard systolic arrays.
-
Exploring Bounded Component Analysis Using an $\ell_\infty$ Norm Criterion
Minimizing the sum of ℓ∞ norms enables separation of antisparse bounded sources via PCA followed by Givens rotations optimization, with claimed superior performance over prior methods in simulations.
-
Interpreting Multi-Branch Anti-Spoofing Architectures: Correlating Internal Strategy with Empirical Performance
A framework using covariance-based spectral signatures and TreeSHAP attributions on AASIST3 branches identifies four operational archetypes and a flawed specialization mode that explains high error rates on specific spoofing attacks.
-
Anomaly Detection for Sparse and Irregular Multivariate Time Series with Latent SDEs
Latent SDE generative model for anomaly detection in sparse irregular multivariate time series outperforms baselines on six benchmarks and stays robust under severe sparsity.
-
Head-Pose-Aware Visual Speech Recognition with FiLM Modulation
HP-VSR-ResFiLM adds a single residual FiLM modulation block conditioned on head pose to a CNN visual encoder, yielding WER of 25.0% on LRS2 and 33.2% on LRS3 under standard training conditions.
-
A Decomposed Retrieval-Edit-Rerank Framework for Chord Generation
The RER framework decomposes chord generation into retrieval, editing, and reranking stages to outperform end-to-end models in balancing stylistic diversity with music-theoretic feasibility.
-
Deep Reinforcement Learning for Spacecraft Attitude Control During Atmospheric Re-Entry
Hybrid RL-PID controllers track angle of attack better and show greater robustness than PID alone within a defined operational envelope for re-entry attitude control.
-
Parameter-Efficient CT Reconstruction via Deep Graph Laplacian Regularization
Deep GLR combines graph Laplacian regularization with three lightweight CNN modules in a proximal optimization framework to reach 30.70 dB PSNR on LoDoPaB-CT using 5.8x fewer parameters and 30x less data per dB gain than typical deep methods.
-
Modality vs. Morphology: A Framework for Time Series Classification for Biological Signals
A review synthesizes evidence from EEG, EMG, ECG, PPG and ocular signals to argue that waveform morphology, rather than modality or model class, primarily determines TSC performance and interpretability.
-
Debunking Grad-ECLIP: A Comprehensive Study on Its Incorrectness and Fundamental Principles for Model Interpretation
Grad-ECLIP is an equivalent but flawed variant of attention-based interpretation, with two principles proposed to ensure model explanations reflect the original model.
-
Comparing Commercial Depth Sensor Accuracy for Medical Applications
Zivid 2M+ 60 outperformed Intel RealSense D405, PMD Flexx2, and Stereolabs ZED 2i on all tested specimens and metrics; ZED ranked second on real tissue but last on the phantom.
-
SoK: A Comprehensive Analysis of the Current Status of Neural Tangent Generalization Attacks with Research Directions
NTGA is the first clean-label generalization attack under black-box settings but is vulnerable to adversarial training and image transformations, with newer attacks outperforming it.