REVIEW 12 cited by
Transformer Neural Processes: Uncertainty-Aware Meta Learning Via Sequence Modeling
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Transformer Neural Processes: Uncertainty-Aware Meta Learning Via Sequence Modeling
read the original abstract
Neural Processes (NPs) are a popular class of approaches for meta-learning. Similar to Gaussian Processes (GPs), NPs define distributions over functions and can estimate uncertainty in their predictions. However, unlike GPs, NPs and their variants suffer from underfitting and often have intractable likelihoods, which limit their applications in sequential decision making. We propose Transformer Neural Processes (TNPs), a new member of the NP family that casts uncertainty-aware meta learning as a sequence modeling problem. We learn TNPs via an autoregressive likelihood-based objective and instantiate it with a novel transformer-based architecture. The model architecture respects the inductive biases inherent to the problem structure, such as invariance to the observed data points and equivariance to the unobserved points. We further investigate knobs within the TNP framework that tradeoff expressivity of the decoding distribution with extra computation. Empirically, we show that TNPs achieve state-of-the-art performance on various benchmark problems, outperforming all previous NP variants on meta regression, image completion, contextual multi-armed bandits, and Bayesian optimization.
Forward citations
Cited by 12 Pith papers
-
Revisiting Neural Processes via Fourier Transform and Volterra Series
Set Fourier convolutions plus a Volterra cascade yield scalable, translation-equivariant CNPs that handle irregular inputs with global receptive fields and beat strong baselines.
-
PRIM: Meta-Learned Bayesian Root Cause Analysis
PRIM meta-learns a Model-Averaged Causal Estimation transformer to perform Bayesian RCA by marginalizing structural uncertainty over synthetic causal priors, achieving 17ms inference on systems up to 100 variables.
-
PRIM: Meta-Learned Bayesian Root Cause Analysis
PRIM is a meta-learned Bayesian RCA method that marginalizes structural uncertainty via a MACE transformer neural process for zero-shot inference on systems up to 100 variables.
-
Expanders Meet Reed-Muller: Easy Instances of Noisy k-XOR
Explicit near-optimal expanders exist for which noisy k-XOR is polynomial-time solvable, falsifying conjectures that expansion implies hardness.
-
Transformer Neural Processes - Kernel Regression
TNP-KR adds a kernel regression transformer block, kernel attention bias, scan attention for translation invariance, and deep kernel attention to achieve lower complexity and state-of-the-art results on meta-regressio...
-
Revisiting Neural Processes via Fourier Transform and Volterra Series
Introduces SFConvCNPs and SFVConvCNPs using set Fourier convolutions and Volterra expansions for translation-equivariant neural processes on irregular data with global receptive fields and linear scaling.
-
Spectral Transformer Neural Processes
STNPs extend TNPs with a spectral aggregator that estimates context spectra, forms spectral mixtures, and injects task-adaptive frequency features to better handle periodicity.
-
PRIM: Meta-Learned Bayesian Root Cause Analysis
PRIM is a meta-learned Bayesian RCA method that identifies root causes by averaging over many possible causal structures via a prior-fitted neural process, achieving 17 ms inference on systems up to 100 variables.
-
Uncertainty-Aware Foundation Models for Clinical Data
The work introduces uncertainty-aware foundation models for clinical data by learning set-valued patient representations that enforce consistency across partial observations and integrate multimodal self-supervised ob...
-
Learning-To-Measure: In-Context Active Feature Acquisition
One transformer, pretrained with sequence modeling, greedily acquires features to reduce prediction uncertainty across multiple tasks with missing data, matching or beating task-specific AFA baselines without per-task...
-
Scalable Spatiotemporal Inference with Biased Scan Attention Transformer Neural Processes
BSA-TNP is a new neural process model with KRBlocks and biased scan attention that claims to match top accuracy while scaling inference to over 1M points in under a minute on a single GPU and supporting translation in...
-
Bayesian Inference of Discretization Error Means in ODEs via Ensemble Kalman Filtering
A Bayesian state-space model with an Ensemble Kalman Filter infers the mean of ODE discretization errors from noisy observations, using a step-size-dependent Markov prior whose convergence is proven.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.