A β-VAE-GAN plus sensor-conditioned Transformer with Easy Attention forecasts near-wall turbulence in the Minimal Flow Unit, recovering 87% turbulent kinetic energy in 4D latent space and maintaining accuracy over 17288 t+ from 128 t+ initialization while reconstructing 82% TKE end-to-end.
Cyclical Annealing Schedule: A Simple Approach to Mitigating KL Vanishing
8 Pith papers cite this work, alongside 169 external citations. Polarity classification is still indexing.
abstract
Variational autoencoders (VAEs) with an auto-regressive decoder have been applied for many natural language processing (NLP) tasks. The VAE objective consists of two terms, (i) reconstruction and (ii) KL regularization, balanced by a weighting hyper-parameter \beta. One notorious training difficulty is that the KL term tends to vanish. In this paper we study scheduling schemes for \beta, and show that KL vanishing is caused by the lack of good latent codes in training the decoder at the beginning of optimization. To remedy this, we propose a cyclical annealing schedule, which repeats the process of increasing \beta multiple times. This new procedure allows the progressive learning of more meaningful latent codes, by leveraging the informative representations of previous cycles as warm re-starts. The effectiveness of cyclical annealing is validated on a broad range of NLP tasks, including language modeling, dialog response generation and unsupervised language pre-training.
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 8roles
background 1polarities
background 1representative citing papers
ACE-LoRA introduces adaptive orthogonal decoupling and rank-invariant compression for continual image editing in diffusion models, plus the CIE-Bench benchmark.
Octopus introduces history-free gradient orthogonalization in a two-stage finetuning framework to achieve state-of-the-art continual learning results for multimodal LLMs on the UCIT benchmark.
VIVALDy is a hybrid β-VAE-GAN plus bidirectional transformer framework that reconstructs and predicts turbulent flow around a one-degree-of-freedom moving cylinder using only cylinder displacement as input.
TabICL scales in-context learning to large tabular data via column-then-row attention for row embeddings followed by a transformer, matching TabPFNv2 speed and performance while outperforming it and CatBoost on datasets over 10K samples.
Ensemble-based method of moments on softmax outputs produces stable Dirichlet predictive distributions that improve uncertainty-guided tasks like selective classification over evidential deep learning.
GCVAE is a variational autoencoder that structures its latent space as a Gaussian mixture and optimizes a variational objective to make the representation maximally informative about a user-chosen guiding variable, enabling context-specific clusters.
NetVAD uses a strictly identifier-free VAE on frozen foundation model embeddings, trained solely on benign traffic, to achieve 98% micro F1 and 96% macro F1 on ToN-IoT for unsupervised intrusion detection.
citing papers explorer
-
A Hybrid Generative Reduced-Order Model for the Minimal Flow Unit
A β-VAE-GAN plus sensor-conditioned Transformer with Easy Attention forecasts near-wall turbulence in the Minimal Flow Unit, recovering 87% turbulent kinetic energy in 4D latent space and maintaining accuracy over 17288 t+ from 128 t+ initialization while reconstructing 82% TKE end-to-end.
-
ACE-LoRA: Adaptive Orthogonal Decoupling for Continual Image Editing
ACE-LoRA introduces adaptive orthogonal decoupling and rank-invariant compression for continual image editing in diffusion models, plus the CIE-Bench benchmark.
-
Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language Models
Octopus introduces history-free gradient orthogonalization in a two-stage finetuning framework to achieve state-of-the-art continual learning results for multimodal LLMs on the UCIT benchmark.
-
VIVALDy: A Hybrid Generative Reduced-Order Model for Turbulent Flows, Applied to Vortex-Induced Vibrations
VIVALDy is a hybrid β-VAE-GAN plus bidirectional transformer framework that reconstructs and predicts turbulent flow around a one-degree-of-freedom moving cylinder using only cylinder displacement as input.
-
TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
TabICL scales in-context learning to large tabular data via column-then-row attention for row embeddings followed by a transformer, matching TabPFNv2 speed and performance while outperforming it and CatBoost on datasets over 10K samples.
-
Ensemble-Based Dirichlet Modeling for Predictive Uncertainty and Selective Classification
Ensemble-based method of moments on softmax outputs produces stable Dirichlet predictive distributions that improve uncertainty-guided tasks like selective classification over evidential deep learning.
-
From Unsupervised to Guided Clustering: A Variational Implementation
GCVAE is a variational autoencoder that structures its latent space as a Gaussian mixture and optimizes a variational objective to make the representation maximally informative about a user-chosen guiding variable, enabling context-specific clusters.
-
NetVAD: Foundation-Model Representation Learning for Identifier-Free Unsupervised Intrusion Detection
NetVAD uses a strictly identifier-free VAE on frozen foundation model embeddings, trained solely on benign traffic, to achieve 98% micro F1 and 96% macro F1 on ToN-IoT for unsupervised intrusion detection.