{"id":"3a172f53-1d7b-4a7b-8433-eccbf37254b3","arxiv_id":"2412.11519","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"LineArt transfers reference-photo appearance onto complex design drawings without training, using multi-frequency line fusion and a two-stage painting process, and introduces the ProLines dataset.","lead":"This paper presents LineArt, a training-free pipeline that transfers a reference photo's materials and textures onto detailed line drawings using ControlNet, IP-Adapter, and Stable Diffusion. It uses multi-frequency line fusion and a two-stage painting process, and introduces a new design-drawing dataset, ProLines.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The soft-edge bootstrap in §4.1 is circular: high-frequency constraints are extracted from G_initial, an image already generated using the appearance reference, so structural fidelity to the original line drawing is not established.","rationale":"The reader's weakest_assumption targets the latent-space brightness formula, which is indeed under-derived and hand-tuned. However, the more load-bearing issue is the circular soft-edge bootstrap: the final ControlNet conditioning includes edges extracted from G_initial, an image already generated using the appearance reference. This directly affects the 'preserving structural accuracy' half of the central claim. Even if the latent shift were perfect, structural fidelity could still be compromised by appearance-derived soft edges. Conversely, if the bootstrap is necessary, the method becomes a two-pass appearance-guided synthesis rather than a structure-preserving transfer. The paper's ablation study reportedly includes 'the role of soft edges,' but it does not test whether soft edges come from G_initial versus the original drawing. A simple replacement test would settle whether the claimed structural control is real. This concern does not change the reader's conditional verdict, so the recommendation remains conditional on the authors providing that ablation and clarifying the metric references in the main text.","tokens_in":16436,"tokens_out":9821,"duration_ms":86866,"concrete_test":"Run the ProLines evaluation with S_soft replaced by Haar(L_original) (or a standard edge map of the input line drawing), keeping all other components fixed. If SSIM/CD or qualitative structural fidelity degrade significantly, the method depends on appearance-derived soft edges and the claimed preservation of the original line drawing is not supported. If performance is unchanged, the bootstrap is unnecessary and should be removed for a cleaner method.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In §4.1, the authors compute Condition_geometry = Fusion(L_double, L_single, S_soft), where S_soft = Haar(G_initial). G_initial is obtained by first generating from the appearance image using ControlNet-depth and the line drawing's predicted depth. Thus the soft edges injected into the final ControlNet conditioning are not derived from the input line drawing but from an image that already contains the reference appearance. This creates a circular dependency: the same appearance image influences the final output both through IP-Adapter embedding and through the geometric conditioning. The paper claims these soft edges 'capture spatial gradients and texture information' and 'act as high-frequency constraints,' but this conflates texture with structure, meaning the final image's edges may be dictated by the reference photo's textures rather than by the original line drawing. That directly threatens the central claim of 'preserving structural accuracy.' No ablation is presented that replaces S_soft with edges computed purely from L_original, so the contribution of the multi-frequency fusion to structure preservation is not isolated from this bootstrap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"LineArt proposes a training-free pipeline for transferring appearance from a reference photo to a detailed line drawing. It builds on Stable Diffusion, ControlNet, and IP-Adapter, and consists of two stages: a multi-frequency line fusion module (single lines, double lines, and soft edges) for structure guidance, and a two-part painting process (Base Layer Shaping via Retinex-based brightness initialization, and Surface Layer Coloring via texture synthesis and selective attention-layer injection). The paper also introduces the ProLines dataset of 5,101 filtered design drawings and reports quantitative, qualitative, and user-study comparisons against several state-of-the-art methods. The central claim is that LineArt outperforms existing methods in accuracy, realism, and material precision while preserving structural details of the input line drawing.","tokens_in":16723,"tokens_out":6215,"duration_ms":61220,"significance":"Training-free appearance transfer for professional line drawings is practically valuable, and the paper's analogy to classical underpainting/glazing is a creative way to structure the pipeline. The proposed ProLines dataset is a potentially useful resource for the sketch-to-image community, and the use of standard components with no training makes the method easy to adopt if the claims hold. However, the current evidence is not yet convincing: the soft-edge construction is circular, the latent brightness initialization is heuristic, key ablations are deferred to an absent supplementary document, and the quantitative evaluation lacks uncertainty estimates. The paper also does not mention code release or dataset availability in the main text. Strengths include a multi-metric comparison, a user study, and a method description detailed enough to be reimplemented; the weaknesses are concentrated in the validation of the central claims.","major_comments":[{"comment":"The soft-edge construction is circular. The paper computes S_soft = Haar(G_initial), where G_initial is itself generated using ControlNet-depth conditioned on both the line drawing and the appearance image. Therefore Condition_geometry = Fusion(L_double, L_single, S_soft) contains high-frequency information derived from the reference appearance, not only from the original line drawing. This conflates texture with structure and directly threatens the central claim of structure preservation. Please ablate S_soft by replacing it with an edge/high-frequency map computed solely from L_original, and report structure metrics (SSIM/CD) for the variant without S_soft and for the appearance-free soft-edge variant. Without this ablation, the contribution of multi-frequency line fusion to structural fidelity is not established.","section":"Section 4.1"},{"comment":"The Base Layer Shaping initialization x'_0 = Lmean + (x0 - Lmean) * 0.5 lacks a derivation or controlled validation. The paper does not explain why a brightness statistic of the pixel-space Retinex illumination can be additively transplanted into the VAE latent x0, nor what Lmean and x0 are precisely (a scalar, a per-channel vector, or a latent tensor), and the factor 0.5 appears hand-picked. The claim that this creates a 'soft link' between the lighting of the reference and the generated image is not supported by theory or by quantitative ablation; Figure 7(b) is only qualitative. Please clarify the notation, provide a derivation or a rigorous empirical study, and report quantitative ablations over Lmean values and the 0.5 factor to justify this key design choice.","section":"Section 4.2"},{"comment":"The quantitative evaluation reports single values without error bars, statistical significance, or the number of images evaluated per dataset. Several comparisons are close or even reverse: for Bronze, DreamBooth achieves higher PSNR (27.15) than LineArt (25.20), and the FID gap between LineArt and the baseline is small in several rows. The claim that LineArt 'performs better in accuracy, realism, and material precision compared to SOTAs' is therefore not statistically substantiated. Please report mean and standard deviation over multiple seeds or test splits, run significance tests, and specify the evaluation set size and protocol for each metric.","section":"Table 1 / Section 6.1"},{"comment":"The ablation study, which addresses four central design choices (double lines, soft edges, surface-layer coloring, and patch size), is only described as being in the supplementary materials; the main text contains no ablation table or figure, and the arXiv version does not appear to include the supplement. The user study details are also deferred. Since these ablations directly support the method's core contributions, the main text should include at least a compact ablation table and the essential user-study protocol (number of participants, number of stimuli, ranking scale, and statistical analysis).","section":"Section 6.3"},{"comment":"The ProLines dataset is filtered to hand-chosen IC intervals (e.g., Bronze 0.2576–0.2903, ImageNet-Sketch 0.2500–0.2650) followed by manual screening by the authors. The paper does not state what fraction of images was removed at each stage, whether annotators agreed, or whether the thresholds were selected before or after observing model performance. This creates a risk that the evaluation is conducted on a self-selected subset favorable to the proposed method. Please provide a sensitivity analysis across IC thresholds and report results on a held-out random sample of the unfiltered source data.","section":"Section 5"}],"minor_comments":[{"comment":"There are several typos and grammatical errors: the section title 'Knowleged-guided Structure Preservation' should be 'Knowledge-guided', 'Accoding' should be 'According', and Figure 7's caption contains 'Origianl' instead of 'Original'.","section":"Abstract and Section 4.2"},{"comment":"The notation is inconsistent: 'Conditiongeometry' appears in Section 4.1 while 'Condition_geometry' appears elsewhere; the operators Fusion, TS, and moveBackground are used without precise definitions. Please define all operators and use consistent notation throughout.","section":"Section 4.1 and Section 4.2"},{"comment":"The metrics GLCM and CH are not defined in the main text. Please specify the exact distance or statistic used for the gray-level co-occurrence matrix and the color histogram loss so that the reported numbers are reproducible.","section":"Section 6.1"},{"comment":"The 'Bronze' dataset source is not cited in the reference list, and no dataset URL or release statement is provided for ProLines. Please add the missing citation and state how the dataset will be distributed.","section":"Section 5"},{"comment":"The list of compared methods is not consistent between the qualitative figures and Table 1: Figure 1 shows T2I-Adapter-SDXL while Table 1 includes DreamBooth and InstructPix2Pix but not T2I-Adapter-SDXL. Please align the comparison methods across all experiments.","section":"Figures 1 and 8"}],"recommendation":"major_revision","confidential_remarks":"This is a borderline paper. The central idea is appealing and relevant, but the validation is not yet convincing. The main scientific risk is the circularity in the soft-edge bootstrap: if S_soft is derived from a first-pass generation already influenced by the reference appearance, the structure-preservation claim is not established. A revision that replaces S_soft with an appearance-free edge map and reports quantitative ablations would substantially strengthen the paper. I also recommend asking the authors to include the ablation and user-study details in the main text, release code and the ProLines dataset, and add uncertainty estimates to the quantitative evaluation. If these points are addressed, I would be willing to see a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"LineArt is a plausible training-free pipeline for transferring materials onto design line drawings, and the new ProLines dataset is the most concrete contribution. The method composes known pieces—ControlNet, IP-Adapter, Retinex, morphological filtering, patch-based texture reassembly—in a way I haven't seen before. The two-stage split into Base Layer Shaping and Surface Layer Coloring is well motivated by the painting analogy and maps to real design decisions. That part reads coherently.\n\nThe weak spot is the soft-edge bootstrap in §4.1. Ssoft = Haar(Ginitial), where Ginitial is a first pass already conditioned on the reference appearance. So the \"high-frequency constraints\" are not derived from the original line drawing; they're derived from an image that already contains the transferred texture. That makes it hard to defend the claim that structure preservation comes from the line drawing rather than from the reference photo. The paper doesn't ablate Ssoft extracted from Loriginal, so the effect of the fusion module on fidelity is not isolated. This needs to be addressed before the central claim holds.\n\nThe other soft spots are evaluation hygiene: no error bars or significance tests, hand-picked ProLines thresholds, ablation study only in supplementary, no code or data release. The latent brightness initialization is a heuristic—shifting latent noise by Lmean and a 0.5 factor—with no derivation from pixel-space illumination. It may work empirically, but it's not justified in the text.\n\nThe citation pattern looks fine; the related work is appropriate and the authors acknowledge limitations honestly. The conclusion explicitly says the method struggles with excessive shadows and complex textures, which is honest but should be reflected in the claims.\n\nWho this is for: people in sketch-to-image or appearance transfer, especially with design drawings in mind. It deserves a serious referee, because the dataset and pipeline are real, but the authors need to release code/data and add the missing ablations and statistics. If they do, this could become a solid subfield contribution. I'd want to see the revised version before citing it in my own work.","headline":"A useful training-free appearance transfer pipeline with a valuable new dataset, but the evaluation hygiene and a circular soft-edge bootstrap keep the central claim from being fully supported.","tokens_in":17202,"tokens_out":2758,"would_cite":false,"duration_ms":27173,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LineArt transfers complex materials and lighting onto detailed design drawings with a training-free diffusion pipeline that outperforms current state-of-the-art methods in accuracy, realism, and material precision.","keywords":["LineArt","appearance transfer","design drawings","diffusion model","training-free","multi-frequency lines fusion","Retinex","ProLines dataset"],"falsifier":"Run a controlled sweep: fix one line drawing and one seed, vary only the reference photo's brightness (or its Retinex-derived $L_{\\text{mean}}$) across a wide range, and measure whether the generated image's mean illumination tracks the reference's $L_{\\text{mean}}$ monotonically. If output brightness is insensitive to $L_{\\text{mean}}$, or is non-monotonic, the core mechanism of Base Layer Shaping is falsified; a complementary check replaces the shifted latent $x'_0$ by the unshifted $x_0$ and compares lighting-aware metrics.","tokens_in":16246,"feed_emoji":"🎨","tokens_out":4078,"duration_ms":35588,"temperature":0.7,"pith_summary":"LineArt is a training-free framework for transferring the appearance of a reference photo — its materials, textures, colors, and lighting — onto a professional line drawing while keeping the drawing's fine structure intact. The paper argues that text prompts and existing image-translation methods fail on detailed design drawings, and that the right way is to guide a pretrained diffusion model with structure cues decomposed into three edge levels and with appearance cues split into a lighting stage and a texture stage. If the claim holds, designers can render technical drawings with realistic materials without 3D modeling, physical property specifications, or paired training data. The paper also introduces ProLines, a professional design drawing dataset, to evaluate the method.","feed_headline":"LineArt copies materials onto design drawings without training","feed_subtitle":"A two-stage diffusion pipeline preserves fine line structure while matching reference lighting and texture.","key_machinery":"The framework's two engines are the multi-frequency lines fusion module and the two-part painting process. The first decomposes the input line drawing into three levels of edge assertions — a single-line outer contour from the mask, double lines from morphological dilation-erosion, and soft edges extracted via Haar filtering of an initial depth-guided generation — which are fused and fed to ControlNet as geometric conditioning. The second splits appearance transfer: Base Layer Shaping computes a multi-scale Retinex illumination image, extracts its brightness mean $L_{\\text{mean}}$, and shifts the initial latent noise by $x'_0 = L_{\\text{mean}} + (x_0 - L_{\\text{mean}}) \\times 0.5$ to pre-bias lighting, while Surface Layer Coloring removes the reference background, reassembles the remaining pixels into a synthetic texture image, and injects its embedding into specific U-net attention layers to control material and color.","core_discovery":"The central discovery is that a knowledge-guided, fully training-free pipeline built on ControlNet, IP-Adapter, and Stable Diffusion can outperform state-of-the-art appearance-transfer and sketch-to-image methods specifically on detailed, professional design drawings. By decomposing the input drawing into single lines, double lines, and soft edges, and by separating the colorization into Base Layer Shaping (driving the initial latent with a brightness analysis of the reference's Retinex illumination) and Surface Layer Coloring (injecting a reassembled texture map into chosen U-net attention layers), the method claims high fidelity to both the drawing's structure and the reference's material, with quantitative gains on eight metrics and user-study wins across edge fidelity, appearance, and overall perception.","pith_inferences":["The pixel-space-to-latent-space brightness transfer used in Base Layer Shaping suggests a testable extension: replacing the hand-set 0.5 factor with a calibrated or per-image adaptive scale could make lighting control more precise and less dependent on the reference photo's dynamic range.","Because the method is assembled from modular components, an ablation that swaps the Retinex brightness analysis for a simple average brightness, or the patch reassembly for plain cropping, would reveal how much of the measured gain comes from the knowledge-guided decomposition versus the base model itself.","The same hierarchical edge decomposition could serve neighboring tasks, such as line-art colorization, sketch-based editing, or material preview in product design, wherever preserving fine structural detail is the bottleneck.","A direct stress test would apply the method to drawings with heavy pre-existing shading or dense cross-hatching, since the paper states these are failure cases and the framework's structure-preservation claims are most at risk there."],"forward_implications":["Designers can preview materials on technical drawings without any training or 3D modeling, needing only a line drawing and a reference photo.","The three-level line decomposition gives control over pattern thickness and texture density, so the same framework adapts to different drawing complexities.","Separating lighting from texture avoids the attribute degradation that appears when a single text prompt or image prompt tries to convey both at once.","The ProLines dataset provides a professional-drawing benchmark for future sketch-to-image and appearance-transfer research.","Because the pipeline is training-free, it can be rebuilt on any future diffusion backbone or adapter with the same structure and appearance guidance modules."],"supporting_citations":[{"why":"ControlNet supplies the conditional structure-injection backbone that keeps the generated image aligned with the drawing's geometry.","marker":"[95]"},{"why":"IP-Adapter provides the appearance embedding mechanism through which the reference image's material and color features enter the diffusion process.","marker":"[93]"},{"why":"Depth Anything predicts depth from the line drawing, enabling the initial generation whose Haar-filtered soft edges become the high-frequency texture guide.","marker":"[91]"},{"why":"Stable Diffusion is the pretrained generative model the whole training-free pipeline builds on.","marker":"[67]"},{"why":"ZeST is the leading zero-shot material-transfer baseline the paper compares against and claims to beat in accuracy and detail preservation.","marker":"[16]"},{"why":"Cross-Image Attention is a primary zero-shot appearance-transfer baseline, used to show the structure-loss failure mode the method is designed to overcome.","marker":"[1]"},{"why":"David Marr's assertion-based theory of visual representation motivates the three-level edge decomposition used for structure preservation.","marker":"[53,54]"}],"fun_headline_variants":["LineArt transfers materials to design drawings without training","Training-free diffusion matches reference lighting on line art","Two-stage pipeline preserves structure in material transfer","Knowledge-guided method beats SOTA on design drawing colorization"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that shifting the initial latent noise by a pixel-domain brightness statistic of the reference photo transfers the reference's lighting to the final image; this mapping is asserted rather than derived, and the 0.5 scale is chosen by hand, so if that linkage is wrong the Base Layer Shaping stage does not do its claimed work.","fun_headline_variants_meta":{"raw":{"variants":["LineArt transfers materials to design drawings without training","Training-free diffusion matches reference lighting on line art","Two-stage pipeline preserves structure in material transfer","Knowledge-guided method beats SOTA on design drawing colorization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000216,"raw_usage":{"total_tokens":1394,"prompt_tokens":871,"completion_tokens":523,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":487,"completion_tokens_details":{"reasoning_tokens":463}},"tokens_in":487,"tokens_out":523,"duration_ms":5258,"temperature":1.0,"reasoning_tokens":463,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:50:35.624511+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled sweep: fix one line drawing and one seed, vary only the reference photo's brightness (or its Retinex-derived $L_{\\text{mean}}$) across a wide range, and measure whether the generated image's mean illumination tracks the reference's $L_{\\text{mean}}$ monotonically. If output brightness is insensitive to $L_{\\text{mean}}$, or is non-monotonic, the core mechanism of Base Layer Shaping is falsified; a complementary check replaces the shifted latent $x'_0$ by the unshifted $x_0$ and compares lighting-aware metrics.","supporting_citations":[{"cited_title":"Adding conditional control to text-to-image diffusion models","cited_arxiv_id":null,"evidence_quote":"ControlNet supplies the conditional structure-injection backbone that keeps the generated image aligned with the drawing's geometry."}],"review_version":1}