Pith. sign in

REVIEW 9 cited by

Masked Audio Generation using a Single Non-Autoregressive Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.04577 v2 pith:YWUCJLFF submitted 2024-01-09 cs.SD cs.AIcs.LGeess.AS

classification cs.SDcs.AIcs.LGeess.AS
keywords magnetautoregressivenon-autoregressiveaudiogenerationmaskedsequencewhile
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce MAGNeT, a masked generative sequence modeling method that operates directly over several streams of audio tokens. Unlike prior work, MAGNeT is comprised of a single-stage, non-autoregressive transformer. During training, we predict spans of masked tokens obtained from a masking scheduler, while during inference we gradually construct the output sequence using several decoding steps. To further enhance the quality of the generated audio, we introduce a novel rescoring method in which, we leverage an external pre-trained model to rescore and rank predictions from MAGNeT, which will be then used for later decoding steps. Lastly, we explore a hybrid version of MAGNeT, in which we fuse between autoregressive and non-autoregressive models to generate the first few seconds in an autoregressive manner while the rest of the sequence is being decoded in parallel. We demonstrate the efficiency of MAGNeT for the task of text-to-music and text-to-audio generation and conduct an extensive empirical evaluation, considering both objective metrics and human studies. The proposed approach is comparable to the evaluated baselines, while being significantly faster (x7 faster than the autoregressive baseline). Through ablation studies and analysis, we shed light on the importance of each of the components comprising MAGNeT, together with pointing to the trade-offs between autoregressive and non-autoregressive modeling, considering latency, throughput, and generation quality. Samples are available on our demo page https://pages.cs.huji.ac.il/adiyoss-lab/MAGNeT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion

    cs.SD 2026-08 conditional novelty 8.0 of 10

    InvFlowFD measures music quality by inverting audio through a flow matching model and computing the distance of the inverted latents to the model's Gaussian prior.

  2. SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation

    eess.AS 2026-08 conditional novelty 7.0 of 10

    SemBridge supervises autoregressive states with discrete semantic tokens during training, improving content fidelity of continuous-latent speech generation without changing inference.

  3. HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion

    cs.CV 2026-08 reject novelty 6.0 of 10

    HarmoniDPO pairs global and frame-level video features with preference-style optimization to generate audio from silent video, reporting improved synchronization and quality metrics over prior V2A baselines.

  4. A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies

    eess.AS 2026-08 conditional novelty 6.0 of 10

    The deciding factors for modeling an audio latent are dependency horizon and conditional ambiguity, viewed jointly with the representation, not whether the latent is discrete or continuous.

  5. Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model

    cs.SD 2026-08 conditional novelty 6.0 of 10

    Diff-Symbo generates long, text-controlled symbolic music by autoregressively extending 8-bar latent diffusion segments conditioned on the previous segment's latent.

  6. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  7. TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models

    cs.SD 2025-09 conditional novelty 6.0 of 10

    TTA-Bench offers a seven-dimension, 2,999-prompt evaluation of ten text-to-audio models with 118,000 human ratings, covering quality, robustness, fairness, bias, and toxicity.

  8. LaViDa: A Large Diffusion Language Model for Multimodal Understanding

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A diffusion-based vision-language model matches several autoregressive baselines on multimodal benchmarks while enabling controllable generation and faster decoding at reduced quality.

  9. Improving Controllability and Editability for Pretrained Text-to-Music Generation Models

    cs.SD 2024-11 conditional novelty 2.0 of 10

    A thesis compilation presenting three complementary approaches to improving editing and control of pretrained text-to-music models, with Instruct-MusicGen demonstrating the strongest stem-level editing results.

Pith tools