REVIEW 8 cited by
Discrete Flow Matching
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Despite Flow Matching and diffusion models having emerged as powerful generative paradigms for continuous variables such as images and videos, their application to high-dimensional discrete data, such as language, is still limited. In this work, we present Discrete Flow Matching, a novel discrete flow paradigm designed specifically for generating discrete data. Discrete Flow Matching offers several key contributions:(i) it works with a general family of probability paths interpolating between source and target distributions; (ii) it allows for a generic formula for sampling from these probability paths using learned posteriors such as the probability denoiser ($x$-prediction) and noise-prediction ($\epsilon$-prediction); (iii) practically, focusing on specific probability paths defined with different schedulers improves generative perplexity compared to previous discrete diffusion and flow models; and (iv) by scaling Discrete Flow Matching models up to 1.7B parameters, we reach 6.7% Pass@1 and 13.4% Pass@10 on HumanEval and 6.7% Pass@1 and 20.6% Pass@10 on 1-shot MBPP coding benchmarks. Our approach is capable of generating high-quality discrete data in a non-autoregressive fashion, significantly closing the gap between autoregressive models and discrete flow models.
Forward citations
Cited by 8 Pith papers
-
Hacking Generative Perplexity: Why Unconditional Text Evaluation Needs Distributional Metrics
Zero-parameter naive samplers achieve state-of-the-art generative perplexity while producing incoherent text, proving the metric is unsound; distributional divergences like MAUVE and energy distance correctly rank the...
-
Coupled Continuous-Discrete Generation for Scene Text Image Super-Resolution
A shared transformer trained with continuous flow matching for images and discrete diffusion for text jointly restores scene text images and reads out their characters, removing the external OCR prior.
-
Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
TarFlowLM models language in a continuous latent space with transformer-based autoregressive normalizing flows, using mixture-CDF and Rosenblatt couplings, and reports competitive NELBO on TEXT8 and OpenWebText.
-
CFMI: Flow Matching for Missing Data Imputation
A conditional flow-matching model trained only on observed portions of data imputes missing entries competitively across 24 tabular and two time-series datasets.
-
Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces
A unified diffusion framework with per-modality noise clocks lets one model generate images, text, and tabular data jointly or conditionally in their native spaces.
-
Corrector Sampling in Language Models
A training and sampling method that lets autoregressive LLMs resample earlier tokens in a small window, improving reasoning and coding benchmark scores by about 10% relative after a 100B-token fine-tuning.
-
Accelerated Sampling from Masked Diffusion Models via Entropy Bounded Unmasking
EB-Sampler dynamically unmasks multiple low-entropy tokens per function evaluation, accelerating masked diffusion model sampling by 2-3x with negligible accuracy loss.
-
SynBridge: Bridging Reaction States via Discrete Flow for Bidirectional Reaction Prediction
A bidirectional discrete flow matching model, SynBridge, predicts reaction products and reactants on graph representations and reports state-of-the-art Top-k accuracy on USPTO-50K, USPTO-MIT, and Pistachio.
Discussion (0). Sign in to comment.