{"work":{"id":"3fd87c40-91ee-403b-9781-58b4e2feb625","openalex_id":null,"doi":null,"arxiv_id":"2508.09192","raw_key":null,"title":"Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing","authors":null,"authors_text":"Diffusion llms can do faster-than-ar inference via discrete diffusion forcing , author=","year":2025,"venue":"cs.LG","abstract":"Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive (AR) LLMs for text generation, with the potential to decode multiple tokens in a single iteration. However, none of the existing open-source dLLMs have achieved superior inference speed over AR LLMs of similar size. This paper breaks this barrier based on a simple and effective strategy named discrete diffusion forcing (D2F). D2F equips dLLMs with two key capabilities: (1) block-wise autoregressive generation to enable KV cache utilization; (2) prediction of following tokens without requiring completion of prior blocks for inter-block parallel decoding. In this way, the vanilla dLLMs are refurbished into an AR-diffusion hybrid paradigm for efficient inference. D2F can be implemented with an asymmetric distillation process based on pre-trained dLLMs. We further propose a pipelined parallel decoding algorithm, which enables a trade-off between efficiency and efficacy. Empirically, D2F dLLMs achieve more than $\\mathbf{2.5\\times}$ inference speed than LLaMA3 and Qwen2.5 on GSM8K. Compared to vanilla dLLMs like LLaDA and Dream, the acceleration can be more than $\\mathbf{50\\times}$ while maintaining comparable output quality. The code is available at https://github.com/zhijie-group/Discrete-Diffusion-Forcing.","external_url":"https://arxiv.org/abs/2508.09192","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-11T03:07:51.302542+00:00","pith_arxiv_id":"2508.09192","created_at":"2026-05-11T00:25:53.963341+00:00","updated_at":"2026-07-11T03:07:51.302542+00:00","title_quality_ok":true,"display_title":"Diffusion llms can do faster-than-ar inference via dis- crete diffusion forcing.arXiv preprint arXiv:2508.09192","render_title":"Diffusion llms can do faster-than-ar inference via dis- crete diffusion forcing.arXiv preprint arXiv:2508.09192"},"hub":{"state":{"work_id":"3fd87c40-91ee-403b-9781-58b4e2feb625","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":15,"external_cited_by_count":null,"distinct_field_count":4,"first_pith_cited_at":"2025-12-10T09:26:18+00:00","last_pith_cited_at":"2026-07-07T01:09:54+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-03T20:10:17.092254+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":3}],"polarity_counts":[{"context_polarity":"background","n":3}],"runs":{},"summary":{},"graph":{},"authors":[]}}