A sequence-graph model using gated modulation of methylation signals by eight handcrafted DNA sequence features achieves 3.149 years MAE on 3707 samples, a 12.8% gain over graph baselines.
Kingma and Jimmy Ba
10 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 10representative citing papers
Fitting logic gates as 4D multilinear polynomials with covariance Jacobian selection matches or beats 16D softmax baselines on seven datasets and remains stable at 12-layer depth where the baseline drops 37 points on CIFAR-10.
A residual-corrected ECM-UDE hybrid model outperforms standalone ECM and LSTM baselines in battery terminal voltage prediction, with the largest gains under temperature and drive-cycle distribution shifts.
Prompts can be split into separate roles for sampling design and recovery modeling in generative compressed sensing, with stable recovery bounds for matched prompts and an explicit penalty for mismatch, validated on Stable Diffusion.
PSD kernel densities give a convex MCAR imputer with closed-form marginals, KL consistency rates that adapt to smoothness, and competitive energy-distance performance on small real tables.
AD-CERT uses logit-level adversarial distillation from a robust teacher combined with IBP to achieve state-of-the-art certified performance on robustness benchmarks, improving over feature-space distillation by up to 5.40 percentage points.
Flexformer learns attention kernels by treating spectral frequencies as trainable parameters in random Fourier feature-based linear attention, with stationary and nonstationary variants that outperform fixed-kernel baselines.
OSDN adds online diagonal preconditioning to the Delta Rule, preserving chunkwise parallelism while proving super-geometric convergence and delivering 32-39% recall gains at 340M-1.3B scales.
PILOT is a policy-informed learned optimizer that conditions its update on gradient agreement to adapt between momentum, normalization, and sign-based rules, reporting highest accuracies on FashionMNIST and CIFAR-10 with CNN and ResNet-18.
LBW-Guard is a bounded autonomous control layer above AdamW that improves stability, reduces perplexity, and speeds up training for Qwen2.5 models under learning-rate stress on WikiText-103.
citing papers explorer
-
Bridging Sequence and Graph Structure for Epigenetic Age Prediction
A sequence-graph model using gated modulation of methylation signals by eight handcrafted DNA sequence features achieves 3.149 years MAE on 3707 samples, a 12.8% gain over graph baselines.
-
Fitting Multilinear Polynomials for Logic Gate Networks
Fitting logic gates as 4D multilinear polynomials with covariance Jacobian selection matches or beats 16D softmax baselines on seven datasets and remains stable at 12-layer depth where the baseline drops 37 points on CIFAR-10.
-
Residual-Corrected Equivalent-Circuit Model with Universal Differential Equations for Robust Battery Voltage Prediction under Operating-Condition Shift
A residual-corrected ECM-UDE hybrid model outperforms standalone ECM and LSTM baselines in battery terminal voltage prediction, with the largest gains under temperature and drive-cycle distribution shifts.
-
Active Learning for Conditional Generative Compressed Sensing
Prompts can be split into separate roles for sampling design and recovery modeling in generative compressed sensing, with stable recovery bounds for matched prompts and an explicit penalty for mismatch, validated on Stable Diffusion.
-
Distributionally Faithful Imputation via Positive Semi-Definite Kernel Density Estimation
PSD kernel densities give a convex MCAR imputer with closed-form marginals, KL consistency rates that adapt to smoothness, and competitive energy-distance performance on small real tables.
-
Improving Certified Robustness via Adversarial Distillation
AD-CERT uses logit-level adversarial distillation from a robust teacher combined with IBP to achieve state-of-the-art certified performance on robustness benchmarks, improving over feature-space distillation by up to 5.40 percentage points.
-
Flexformer: Flexible Linear Transformer with Learnable Attention Kernel
Flexformer learns attention kernels by treating spectral frequencies as trainable parameters in random Fourier feature-based linear attention, with stationary and nonstationary variants that outperform fixed-kernel baselines.
-
OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention
OSDN adds online diagonal preconditioning to the Delta Rule, preserving chunkwise parallelism while proving super-geometric convergence and delivering 32-39% recall gains at 340M-1.3B scales.
-
PILOT: Policy-Informed Learned Optimization for Adaptive Deep Network Training
PILOT is a policy-informed learned optimizer that conditions its update on gradient agreement to adapt between momentum, normalization, and sign-based rules, reporting highest accuracies on FashionMNIST and CIFAR-10 with CNN and ResNet-18.
-
Learn-by-Wire Training Control Governance: Bounded Autonomous Training Under Stress for Stability and Efficiency
LBW-Guard is a bounded autonomous control layer above AdamW that improves stability, reduces perplexity, and speeds up training for Qwen2.5 models under learning-rate stress on WikiText-103.