Pith. sign in

REVIEW 9 cited by

Autoregressive Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.02037 v2 pith:BSW7MT2X submitted 2021-10-05 cs.LG stat.ML

classification cs.LGstat.ML
keywords ardmsdiffusionmodelsautoregressivegenerationmodelcompressiondata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Autoregressive Diffusion Models (ARDMs), a model class encompassing and generalizing order-agnostic autoregressive models (Uria et al., 2014) and absorbing discrete diffusion (Austin et al., 2021), which we show are special cases of ARDMs under mild assumptions. ARDMs are simple to implement and easy to train. Unlike standard ARMs, they do not require causal masking of model representations, and can be trained using an efficient objective similar to modern probabilistic diffusion models that scales favourably to highly-dimensional data. At test time, ARDMs support parallel generation which can be adapted to fit any given generation budget. We find that ARDMs require significantly fewer steps than discrete diffusion models to attain the same performance. Finally, we apply ARDMs to lossless compression, and show that they are uniquely suited to this task. Contrary to existing approaches based on bits-back coding, ARDMs obtain compelling results not only on complete datasets, but also on compressing single data points. Moreover, this can be done using a modest number of network calls for (de)compression due to the model's adaptable parallel generation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. S2D2: Fast Decoding for Diffusion LLMs via Training-Free Self-Speculation

    cs.CL 2026-03 conditional novelty 6.5 of 10

    Training-free self-speculation reuses a block-diffusion model’s block-size-1 mode as a local AR verifier, improving accuracy–speed tradeoffs over confidence-threshold decoding.

  2. STRIDE: Post-Training LLMs to Reason and Refine Bio-Sequences via Edit Trajectories

    cs.CE 2026-03 conditional novelty 6.0 of 10

    Training LLMs to emit executable edit trajectories (INSERT/DELETE/REPLACE) from Levenshtein alignments plus policy optimization improves oracle-scored bio-sequence optimization success and novelty.

  3. OmniCache: A Trajectory-Oriented Global Perspective on Training-Free Cache Reuse for Diffusion Transformer Models

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A training-free cache-reuse scheme that spreads computation across the full diffusion trajectory and subtracts estimated noise, accelerating DiT sampling with claimed competitive quality.

  4. PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention

    cs.AI 2026-07 conditional novelty 5.0 of 10

    A hybrid causal-bidirectional attention mask improves adapting pretrained AR models to masked diffusion (WikiText-103 PPL 34.1→28.7) and composes with objective adaptation, though matched AR fine-tuning remains stronger.

  5. Agents of Diffusion: Enhancing Diffusion Language Models with Multi-Agent Reinforcement Learning for Structured Data Generation (Extended Version)

    cs.MA 2026-01 reject novelty 5.0 of 10

    AoD pairs a frozen diffusion language model with two LLM agents that iteratively rewrite prompts from natural-language feedback, reporting better JSON diversity and validity, though the claimed RL mechanism and theore...

  6. Towards High-Fidelity and Controllable Bioacoustic Generation via Enhanced Diffusion Learning

    cs.SD 2025-08 reject novelty 5.0 of 10

    BirdDiff combines a multi-band enhancement stage with a DiffWave-based diffusion generator and multimodal conditioning, reporting substantially better bird-call synthesis metrics than DiffWave on a 12-species propriet...

  7. LVPNet: A Latent-variable-based Prediction-driven End-to-end Framework for Lossless Compression of Medical Images

    eess.IV 2025-06 conditional novelty 5.0 of 10

    LVPNet reports lower bits-per-pixel than prior learned lossless codecs by conditioning pixel predictions on a global multi-scale latent variable with a quantization compensation module.

  8. Solving Inverse Problems via Diffusion-Based Priors: An Approximation-Free Ensemble Sampling Approach

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A weighted-particle sampler evolves the posterior through the diffusion model's reverse dynamics, with theoretical error bounds and improved image reconstructions.

  9. Synthetic Tabular Data Generation: A Comparative Survey for Modern Techniques

    cs.LG 2025-07 conditional novelty 3.0 of 10

    A survey that categorizes tabular data synthesis by generation objectives and adds a benchmark comparison of six models on Adult and CreditRisk.

Pith tools