Stochastic Attention adds calibrated uncertainty to transformer foundation models through inference-time multinomial sampling of attention weights and univariate post-hoc tuning of a concentration parameter.
Title resolution pending
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
MAGELLAN augments LLM agents with online metacognitive LP prediction via semantic generalization to scale curriculum learning in open-ended goal spaces.
DoRA improves LoRA by decomposing weights into magnitude and direction and updating only direction with low-rank matrices, closing much of the gap to full fine-tuning.
citing papers explorer
-
Calibrating Scientific Foundation Models with Inference-Time Stochastic Attention
Stochastic Attention adds calibrated uncertainty to transformer foundation models through inference-time multinomial sampling of attention weights and univariate post-hoc tuning of a concentration parameter.
-
MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces
MAGELLAN augments LLM agents with online metacognitive LP prediction via semantic generalization to scale curriculum learning in open-ended goal spaces.
-
DoRA: Weight-Decomposed Low-Rank Adaptation
DoRA improves LoRA by decomposing weights into magnitude and direction and updating only direction with low-rank matrices, closing much of the gap to full fine-tuning.