Pith. sign in

REVIEW 1 cited by

The DEformer: An Order-Agnostic Distribution Estimating Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.06989 v2 pith:JVNSEV74 submitted 2021-06-13 cs.LG

classification cs.LG
keywords distributionfeatureinputorder-agnosticarchitecturesautoregressivedeformerestimating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Order-agnostic autoregressive distribution (density) estimation (OADE), i.e., autoregressive distribution estimation where the features can occur in an arbitrary order, is a challenging problem in generative machine learning. Prior work on OADE has encoded feature identity by assigning each feature to a distinct fixed position in an input vector. As a result, architectures built for these inputs must strategically mask either the input or model weights to learn the various conditional distributions necessary for inferring the full joint distribution of the dataset in an order-agnostic way. In this paper, we propose an alternative approach for encoding feature identities, where each feature's identity is included alongside its value in the input. This feature identity encoding strategy allows neural architectures designed for sequential data to be applied to the OADE task without modification. As a proof of concept, we show that a Transformer trained on this input (which we refer to as "the DEformer", i.e., the distribution estimating Transformer) can effectively model binarized-MNIST, approaching the performance of fixed-order autoregressive distribution estimating algorithms while still being entirely order-agnostic. Additionally, we find that the DEformer surpasses the performance of recent flow-based architectures when modeling a tabular dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ORIGAMI: A generative transformer architecture for predictions from semi-structured data

    cs.LG 2024-12 conditional novelty 6.0 of 10

    ORIGAMI is a generative transformer with structure-preserving tokenization, key/value position encoding, and grammar-constrained decoding that matches or beats baselines on tabular, multi-label, and code-classification tasks.

Pith tools