{"id":"130f079f-2529-401d-80c4-51ab7e502bc9","arxiv_id":"2601.08303","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A compact elastic diffusion transformer with adaptive sparse attention and knowledge-guided distribution-matching distillation achieves 4-step 1K image generation on a phone in roughly 1.8 seconds.","lead":"SnapGen++ builds a 0.4-billion-parameter diffusion transformer that generates 1024×1024 images in about 1.8 seconds on an iPhone, using sparse attention, elastic sub-network training, and few-step distillation. It is a practical engineering system combining existing ideas to bring transformer-level text-to-image quality to mobile hardware.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2's baselines are marked OOM at 1024×1024 yet still have benchmark scores; the 'surpasses 20× larger' comparison may rest on mismatched evaluation resolutions.","rationale":"The reader's weakest_assumption—that ImageNet validation loss at 256×256 is a reliable proxy for 1024×1024 T2I quality—is reasonable, but my reading identifies a more concrete and immediately checkable weakness: Table 2's internal OOM inconsistency. The central claim is comparative, so the controlledness of the baseline evaluation is load-bearing. If the † baselines were scored at lower resolution, the headline comparison is invalid; the marginal edge over SD3.5-Large makes this especially consequential. The reader did flag the OOM inconsistency in their rationale, but did not treat it as the primary load-bearing assumption, hence partial agreement. This is an addressable issue—release eval configs or rerun at matched resolution—so it strengthens the case for CONDITIONAL rather than moving to ACCEPT or REJECT. The on-device latency claim also omits any text-encoder cost despite use of Gemma3-4b-it in the T2I configuration, which is a secondary concern that reinforces the need for conditionality. Overall, the paper's engineering plausibility and human-eval results support conditional acceptance, but the table inconsistency must be resolved.","tokens_in":34868,"tokens_out":14494,"duration_ms":127796,"concrete_test":"Reproduce the †-marked baseline rows of Table 2 at 1024×1024 using released evaluation code, the same sampler, CFG, and step count as Ours-small (or explicitly identify the resolution/protocol used to obtain the table's baseline scores). If any DPG/GenEval/T2I-CompBench entry changes by more than ~1 point or ~0.01 relative to the table, the 'surpasses up to 20× larger' claim is not established at matched resolution.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the 0.4B model 'surpasses models up to 20× larger' rests on Table 2. The table's footnote states that '† indicates out-of-memory (OOM) at 1024×1024 resolution,' and most baseline rows—PixArt-Σ, SANA, SD3-Medium, SD3.5-Large, Flux.1-dev, etc.—carry that mark. Yet the same rows report DPG-Bench, GenEval, T2I-CompBench, and CLIP scores, and some also report A100 FPS. If those baselines OOM at the stated resolution, the reported numbers must have been produced at a different resolution or with a different protocol (e.g., 512×512, no CFG, smaller batch). T2I quality metrics are resolution-sensitive, so this makes the headline comparison uncontrolled. The concern is concrete: Ours-small's advantage over the closest 20× baseline, SD3.5-Large (8.1B), is within 0.4 points on DPG and 0.01 on GenEval and T2I-CompBench (85.2 vs 85.6, 0.70 vs 0.71, 0.506 vs 0.507). A small evaluation mismatch could flip the claimed superiority. The paper does not state how the OOM rows were evaluated. This is an internal inconsistency rather than a disagreement with consensus, and it needs to be resolved before the central comparative claim is accepted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SnapGen++, a family of efficient diffusion transformers for mobile/edge deployment. Three contributions are claimed: (i) a three-stage DiT architecture with Adaptive Sparse Self-Attention (ASSA), combining compressed global attention and blockwise neighborhood attention; (ii) an elastic training framework that jointly optimizes subnetworks of different widths from one supernetwork; and (iii) Knowledge-guided Distribution Matching Distillation (K-DMD), a step-distillation pipeline that adds a few-step teacher to the DMD objective. The headline claim is that a 0.4B, 4-step variant 'surpasses models up to 20× larger' while running in about 1.8 s on an iPhone 16 Pro Max, and that a 1.6B full variant approaches server-level T2I quality. Experiments include ImageNet-1K validation-loss ablations, T2I benchmarks (DPG-Bench, GenEval, T2I-CompBench, CLIP-Score), latency measurements, a human preference study, and qualitative comparisons.","tokens_in":35298,"tokens_out":4726,"duration_ms":49166,"significance":"The paper addresses a practically important problem — high-fidelity image generation on edge devices — and the proposed components are well-motivated and plausible. The architectural ablations are quantified, and the reported on-device latencies are concrete and specific. If the central comparative claim is correct, the work would be a meaningful step toward making DiT-based T2I generation practical on phones. The elastic training framework and the few-step distillation pipeline also offer useful recipes beyond the specific model. However, the headline comparison is currently undermined by an evaluation-protocol inconsistency in Table 2, and the architecture ablations rest on a single proxy metric whose correlation with perceptual quality is asserted rather than demonstrated. The paper does not release code or evaluation scripts, which would substantially increase confidence in the numerical claims.","major_comments":[{"comment":"The footnote to Table 2 states that '† indicates out-of-memory (OOM) at 1024×1024 resolution,' yet most baseline rows carrying that mark (PixArt-Σ, SANA, SD3-Medium, SD3.5-Large, Flux.1-dev, etc.) report DPG-Bench, GenEval, T2I-CompBench, and CLIP scores. If those models cannot run at 1024×1024 in the stated measurement setup, the numbers must come from a different evaluation protocol (e.g., lower resolution, different sampling steps/CFG, or original papers). T2I benchmarks are resolution-sensitive, so mixing protocols makes the comparison uncontrolled. This matters directly for the central claim 'surpasses models up to 20× larger': the margin over SD3.5-Large is 0.4 DPG points, 0.01 GenEval, and 0.001 T2I-CompBench, so a small protocol mismatch could flip the result. Please specify exactly how each OOM row was evaluated, or separate the 'as-measured-by-us' scores from literature scores","section":"Table 2 and Sec. 4.2"},{"comment":"Every architecture decision — ASSA, three-stage layout, FFN expansion, layer redistribution, GQA — is selected using ImageNet-1K validation loss at 256×256 plus iPhone latency. The paper asserts that validation loss 'shows stronger correlation with perceptual quality and human preference than conventional image metrics such as FID,' but no evidence is provided for this correlation in the context of architecture ablations, and the cited finding [18] does not directly validate this proxy for the architectural variants under consideration. If the proxy is unreliable, the claimed support for the architectural choices collapses. Please provide either (a) FID or human-preference measurements for the ablated variants, or (b) a defensible citation or analysis that validation-loss differences of the observed magnitude (e.g., 0.5130 vs. 0.5090) are perceptually meaningful.","section":"Sec. 3.1, Fig. 3"},{"comment":"The main paper's headline table omits Qwen-Image (20B) and HiDream-I1 (17B), even though both appear in the supplementary detailed tables. Since Qwen-Image is the actual KD teacher and is itself a '20× larger' model, the claim that the 0.4B variant 'surpasses models up to 20× larger' is not tested against the model that defines the teacher quality ceiling. The supplementary numbers show the 0.4B model is far below Qwen-Image on DPG (85.2 vs. 88.3), so the current wording overstates the result relative to the largest relevant baseline. Please include Qwen-Image (and ideally HiDream-I1) in the main comparison table, or qualify the claim to the specific baselines listed.","section":"Table 2 vs. supplementary Tables 2–4"}],"minor_comments":[{"comment":"The K-DMD objective is not fully specified. In Eq. (9), the mapping F and the sampling distribution of τ are not defined; Eq. (10) reintroduces ξ' but the definitions of L^{ξ'}_out and L^{ξ'}_feat are only inferred from Eqs. (6)–(7); and the timestep-aware scaling operator S in Eq. (8) is never given explicitly. Please complete the notation so the method is reproducible.","section":"Sec. 3.3, Eqs. (9)–(10)"},{"comment":"The caption says 'Comparison between Standalone and Elastic training for 0.4B and 2B models,' but the table columns are '0.4B' and '1.6B'. Please align the text and the table.","section":"Table 1"},{"comment":"The main paper says the total on-device runtime is 'around 1.7 s' (Sec. 4.1), while Figure 1 and the supplementary latency table report 1.8 s for the 0.4B model. This inconsistency should be resolved.","section":"Sec. 4.1 and supplementary Sec. A"},{"comment":"The human preference study reports percentages but no error bars, number of participants, or statistical significance tests. Given that some reported differences are small, please add confidence intervals or at least the number of ratings.","section":"Fig. 7"},{"comment":"The paper does not state whether evaluation code or model weights will be released. Given that the main comparison depends on exact evaluation protocols (and the OOM issue above), releasing the evaluation harness would greatly improve trust in the numbers.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the Table 2 protocol inconsistency: if the OOM rows were scored by the authors at a different resolution or sampling setting, the headline 'surpasses 20×' claim is not yet supported. The rest of the paper is a competent systems-engineering contribution, and the architecture/diagnostic story is interesting, but the authors need to clarify the evaluation protocol, add the teacher to the comparison, and defend the validation-loss proxy before the central claims can be accepted. I would be willing to re-review a revision that addresses these points."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the three pieces—ASSA, elastic width slicing, and K-DMD—are coherent, and the on-device 1.8s demo is credible. But the central comparative claim doesn't hold up against the paper's own Table 2. The stress-test note is right: the OOM footnote and the reported baseline scores don't fit together, and the margins versus SD3.5-Large (the closest 20× model) are actually slightly in the baseline's favor on DPG, GenEval, and T2I-CompBench. So \"surpasses models up to 20× larger\" is over-claiming.\n\nWhat's genuinely new: no prior work puts a DiT with adaptive global-local sparse attention, width-sliced elastic training, and few-step distillation onto a phone. The architecture ablations are quantified with validation loss and latency, and the human preference study is a plus. K-DMD—using a few-step LoRA teacher to stabilize DMD for small models—is a sensible engineering contribution.\n\nSoft spots, in order of seriousness. First, the OOM inconsistency: the footnote marks most baselines as OOM at 1024×1024, yet those same rows report DPG, GenEval, T2I-CompBench, and CLIP scores. The paper never states how those baseline numbers were produced. If they were generated at 512 while the proposed models run at 1024, the comparison is uncontrolled. Second, the headline claim is numerically weaker than advertised: Ours-small (0.4B) is essentially tied with SD3.5-Large (8.1B) and in fact slightly behind on three of the four metrics. That doesn't kill the paper, but the authors need to soften or properly qualify the claim. Third, there is no code, no weights, no eval scripts, and no error bars. This is an industry paper, so release may be restricted, but reproducibility is thin. Fourth, the architecture is tuned on ImageNet validation loss at 256 resolution; that's a reasonable proxy, but the transfer to 1024 T2I quality is assumed, not demonstrated. Finally, omitting Qwen-Image from the main table is conspicuous since it is the teacher model.\n\nBottom line: this is a serious engineering paper that deserves peer review, but the evaluation section needs an honest rewrite and the comparative claims need to be scaled back to what the numbers actually show. I'd send it to reviewers with instructions to focus on the resolution/protocol question and to ask for eval code.","headline":"Solid on-device DiT systems paper; the 'surpasses 20×' claim overstates its own Table 2.","tokens_in":35816,"tokens_out":2971,"would_cite":true,"duration_ms":30260,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 0.4B diffusion transformer matches 20x larger models and runs on a phone","keywords":["diffusion transformers","on-device image generation","sparse attention","elastic networks","step distillation","mobile deployment","text-to-image"],"falsifier":"Run a controlled user study with the same prompts used in the paper, pitting the 0.4B on-device model against a server-scale model (e.g., 12B) and check whether the claimed win in realism, fidelity, or alignment holds; alternatively, evaluate the 4-step distilled 0.4B model on a larger, more diverse prompt set and see whether its GenEval and DPG scores collapse.","tokens_in":34818,"feed_emoji":"📱","tokens_out":2558,"duration_ms":24676,"temperature":0.7,"pith_summary":"The paper claims that transformer-based diffusion models, not just U-Nets, can be made to run efficiently on phones without losing the quality advantage transformers bring. It proposes a compact three-stage DiT with adaptive sparse attention, an elastic training scheme that lets one model serve several hardware tiers, and a step-distillation method that cuts sampling to four steps. If the claims hold, a 0.4-billion-parameter model can generate 1024-square images in about 1.8 seconds on a phone, matching or beating models up to twenty times larger on standard text-to-image benchmarks.","feed_headline":"0.4B DiT matches 20x larger models and runs on a phone","feed_subtitle":"Three design moves — sparse attention, elastic training, and 4-step distillation — bring transformer-quality image generation to edge device","key_machinery":"The load-bearing object is the Adaptive Sparse Self-Attention (ASSA) layer, which splits attention into a coarse global branch (key/value features compressed by a strided 2×2 convolution) and a fine local branch (blockwise neighborhood attention with a small number of blocks and radius), adaptively interpolating between the two per head. Around it sits a three-stage transformer (down/middle/up) with token downsampling in the middle, elastic width-sliced supernetwork training, and K-DMD, a distillation objective that adds output-level and feature-level supervision from a few-step teacher to the standard DMD loss.","core_discovery":"The paper's central claim is that a 0.4B-parameter diffusion transformer, running entirely on a mobile device, can match the generation quality of server-scale models with up to 20× more parameters. This is achieved by combining three components: an architecture that replaces full self-attention with an adaptive mix of compressed global attention and blockwise neighborhood attention; an elastic training framework that shares weights across sub-networks of different widths; and a knowledge-guided distribution matching distillation that compresses the sampling process to four steps. The authors support this with benchmark scores (DPG-Bench, GenEval, T2I-CompBench, CLIP), a user study, and on-d","pith_inferences":["If the validation-loss proxy transfers to human preference, the ablation-driven design choices (ASSA, layer distribution, FFN expansion) likely survive scale-up to larger datasets and resolutions; however, that transfer is not demonstrated here.","The comparison to server models relies on benchmarks at 1024 resolution; a direct side-by-side human study with more prompts and diverse devices would show how far the quality claim generalizes.","The elastic framework could be extended to also slice depth and attention heads, not just width, potentially yielding even finer cost-quality trade-offs.","Because the K-DMD teacher is already a few-step model, the method may extend to future one-step or few-step models, possibly enabling real-time generation with fewer than four steps."],"forward_implications":["On-device generation at this quality and latency would make text-to-image features practical in consumer apps without cloud round-trips.","The elastic training result implies one trained model can serve phones, tablets, and servers, replacing per-device fine-tuning.","The 4-step K-DMD result suggests few-step distillation can be applied successfully to small models, not just large ones.","The architectural ablations indicate that sparse attention with global and local branches can recover most of the quality of full attention while cutting latency and memory."],"fun_headline_variants":["Phone DiT: 0.4B matches 20x bigger","0.4B diffusion transformer rivals 20x larger on phone","Edge DiT: 4-step, sparse attention, phone-ready","Diffusion on device: 0.4B DiT outruns 20x models","Mobile DiT: sparse attention + 4-step distillation beats 20x"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The architecture is tuned against ImageNet validation loss at 256 resolution as a proxy for perceptual quality and human preference, and the authors assume this proxy is reliable and that the conclusions transfer to 1024-resolution text-to-image generation.","fun_headline_variants_meta":{"raw":{"variants":["Phone DiT: 0.4B matches 20x bigger","0.4B diffusion transformer rivals 20x larger on phone","Edge DiT: 4-step, sparse attention, phone-ready","Diffusion on device: 0.4B DiT outruns 20x models","Mobile DiT: sparse attention + 4-step distillation beats 20x"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000434,"raw_usage":{"total_tokens":2039,"prompt_tokens":726,"completion_tokens":1313,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":1227}},"tokens_in":470,"tokens_out":1313,"duration_ms":9201,"temperature":1.0,"reasoning_tokens":1227,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T10:51:26.638716+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled user study with the same prompts used in the paper, pitting the 0.4B on-device model against a server-scale model (e.g., 12B) and check whether the claimed win in realism, fidelity, or alignment holds; alternatively, evaluate the 4-step distilled 0.4B model on a larger, more diverse prompt set and see whether its GenEval and DPG scores collapse.","supporting_citations":[],"review_version":1}