Derives inequalities between L1 density distances and mixing-measure discrepancies to obtain posterior contraction rates for Dirichlet process mixtures with unknown shared scale.
Attention is All you Need , volume =
7 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Model collapse occurs in structured interactive learning if and only if the directed interaction graph satisfies a specific topological condition, with finite-sample guarantees for linear regression and asymptotic results for M-estimators.
Bayesian softmax-gated mixture-of-experts models achieve posterior contraction for density estimation and parameter recovery using Voronoi losses, plus two strategies for choosing the number of experts.
Temperature scaling of density-matrix eigenvalues from LLM semantic embeddings optimizes proper-score calibration and corrects systematic overconfidence so entropy equals risk.
ICDN is a neural network that models log-demand from log-prices so elasticities can be derived exactly by differentiation, showing better out-of-sample performance than log-log benchmarks on beer sales data.
Higher-order Langevin dynamics reduce memorization in diffusion models by governing data trajectories with a low-pass-filtered score whose smoothness increases with order.
Adapting autoregressive models via continual pre-training yields diffusion language models from 127M to 7B parameters that outperform prior diffusion models and compete with their autoregressive counterparts on language, reasoning, and commonsense benchmarks.
citing papers explorer
-
Convergence Rates for Latent Mixing Measures in Infinite Homoscedastic Location-Scale Mixture Models
Derives inequalities between L1 density distances and mixing-measure discrepancies to obtain posterior contraction rates for Dirichlet process mixtures with unknown shared scale.
-
When Does Model Collapse Occur in Structured Interactive Learning?
Model collapse occurs in structured interactive learning if and only if the directed interaction graph satisfies a specific topological condition, with finite-sample guarantees for linear regression and asymptotic results for M-estimators.
-
On Bayesian Softmax-Gated Mixture-of-Experts Models
Bayesian softmax-gated mixture-of-experts models achieve posterior contraction for density estimation and parameter recovery using Voronoi losses, plus two strategies for choosing the number of experts.
-
Eigenvalue Calibration for Semantic Embeddings of Large Language Models
Temperature scaling of density-matrix eigenvalues from LLM semantic embeddings optimizes proper-score calibration and corrects systematic overconfidence so entropy equals risk.
-
Integrable Elasticity via Neural Demand Potentials
ICDN is a neural network that models log-demand from log-prices so elasticities can be derived exactly by differentiation, showing better out-of-sample performance than log-log benchmarks on beer sales data.
-
Reducing Diffusion Model Memorization with Higher Order Langevin Dynamics
Higher-order Langevin dynamics reduce memorization in diffusion models by governing data trajectories with a low-pass-filtered score whose smoothness increases with order.
-
Scaling Diffusion Language Models via Adaptation from Autoregressive Models
Adapting autoregressive models via continual pre-training yields diffusion language models from 127M to 7B parameters that outperform prior diffusion models and compete with their autoregressive counterparts on language, reasoning, and commonsense benchmarks.