REVIEW 9 cited by
DINOISER: Diffused Conditional Sequence Learning by Manipulating Noises
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
While diffusion models have achieved great success in generating continuous signals such as images and audio, it remains elusive for diffusion models in learning discrete sequence data like natural languages. Although recent advances circumvent this challenge of discreteness by embedding discrete tokens as continuous surrogates, they still fall short of satisfactory generation quality. To understand this, we first dive deep into the denoised training protocol of diffusion-based sequence generative models and determine their three severe problems, i.e., 1) failing to learn, 2) lack of scalability, and 3) neglecting source conditions. We argue that these problems can be boiled down to the pitfall of the not completely eliminated discreteness in the embedding space, and the scale of noises is decisive herein. In this paper, we introduce DINOISER to facilitate diffusion models for sequence generation by manipulating noises. We propose to adaptively determine the range of sampled noise scales for counter-discreteness training; and encourage the proposed diffused sequence learner to leverage source conditions with amplified noise scales during inference. Experiments show that DINOISER enables consistent improvement over the baselines of previous diffusion-based sequence generative models on several conditional sequence modeling benchmarks thanks to both effective training and inference strategies. Analyses further verify that DINOISER can make better use of source conditions to govern its generative process.
Forward citations
Cited by 9 Pith papers
-
TFG-Flow: Training-free Guidance in Multimodal Generative Flow
TFG-Flow guides multimodal flow models at inference time by weighted Monte Carlo sampling for discrete atom types and gradient ascent for continuous coordinates, improving targeted molecular generation without extra training.
-
Token Time Continuous Diffusion for Language Modeling
A continuous diffusion language model where each token denoises at its own rate—sure tokens first—improves few-step generation over discrete samplers and roughly matches global-time continuous models.
-
LLaDA-VLA: Vision Language Diffusion Action Models
LLaDA-VLA applies a masked diffusion vision-language model to robot control with localized action-token classification and hierarchical decoding, achieving SOTA success rates on SimplerEnv, CALVIN, and real-robot tasks.
-
dKV-Cache: The Cache for Diffusion Language Models
dKV-Cache reuses cached key and value states of decoded tokens during diffusion LM denoising, delivering 2-10x faster inference with near-lossless quality on several benchmarks.
-
DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak
A diffusion-based prompt rewriter that pushes rewritten prompts toward harmless regions of a target model's hidden states achieves higher jailbreak success than existing suffix and template attacks.
-
DiffSLT: Enhancing Diversity in Sign Language Translation via Diffusion Model
DiffSLT uses a latent diffusion model conditioned on fused multi-level visual features to produce diverse, accurate sign language translations, and DiffSLT-P conditions on pseudo-glosses to improve accuracy further.
-
DLM-One: Diffusion Language Models for One-Step Sequence Generation
DLM-One distills a continuous diffusion language model into a one-step student, achieving roughly 500x inference speedup while staying within a few percent of the teacher on BLEU, ROUGE, and BERTScore, with substantia...
-
The Philosophy and Physics of Duality
A philosophical monograph that surveys dualities across physics and proposes a 'geometric view of theories' for theoretical equivalence, realism, and explanation.
-
Diffusion Decoding for Peptide De Novo Sequencing
A diffusion decoder with DINOISER loss raises amino acid recall from 0.081 to 0.454 in Casanovo, while peptide precision and coverage stay at 0 and predicted sequences are much longer than the true peptides.
Discussion (0). Continue with ORCID to comment.