{"id":"a89761fa-cc6a-4e42-a0c0-c33167df4448","arxiv_id":"2504.12605","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Using DeQAScore-derived quality maps to scale diffusion prompt complexity per region improves restoration PSNR, by roughly 1.4 dB over OneRestore on CDD-11 and less than 0.1 dB over T3-DiffWeather on most weather benchmarks.","lead":"AdaQual-Diff restores degraded photos with a diffusion model whose text prompts are tuned per image region, using a quality-scoring AI to work harder on badly damaged areas and leave clean areas alone. Readers in computer vision will find a clean mechanism with modest measured gains, and a case study in how quality metrics can drive generative control.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported CDD-11 result is internally inconsistent (30.11 vs 31.02 dB) and unreproducible without code; the central SOTA claim rests on an unverifiable number.","rationale":"The reader's CONDITIONAL verdict is sound: the ablations show a real effect (Table 3: adaptive prompting 30.11 vs fixed C=30 at 29.87; Table 4: DeQAScore improves over Q-Align and Compare2Score), and the supplementary limitations section is honest about computational costs. I do not dispute that a quality-adaptive prompting mechanism can help. My stress-test focuses on the headline number itself: Section 5.2 claims 31.02/0.9091 while Table 1 reports 30.11/0.9001. Because no code or checkpoint is released, one cannot tell which is genuine; if the true number is 30.11, the text overstates by 0.91 dB, and if it is 31.02, Table 1 is wrong. Either way, the central quantitative claim is not yet reproducible. This is an internal-consistency and reproducibility concern, not a disagreement with consensus or an attack on the idea. A single reproduction test would settle it. I did not select the scalar-mean sufficiency concern as the primary stress test because even if quality-based prompts are not perfectly degradation-type-discriminative, the empirical result could still hold; the numeric discrepancy directly threatens the reported result. The reader's rationale did flag the 31.02/30.11 mismatch among several blocking issues, but the reader's stated weakest assumption was scalar sufficiency, hence partial agreement. The verdict stays CONDITIONAL because the issue is addressable by releasing code and correcting the reported number; it does not require rejecting the underlying approach outright.","tokens_in":20209,"tokens_out":7630,"duration_ms":80474,"concrete_test":"Release the trained checkpoint, inference script, and exact CDD-11 test preprocessing described in Section 5.1, then recompute mean PSNR/SSIM over the official 2,200 test images with a fixed seed. If the reproduced value does not match both 30.11/0.9001 and 31.02/0.9091, the authors must state which is the correct result and revise the abstract and Section 5.2 accordingly; the accepted number should be the one that reproduces from the released artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claim is the CDD-11 result. The text in Section 5.2 states AdaQual-Diff outperforms the previous best method 'by a large margin with 31.02 dB PSNR and 0.9091 SSIM,' but Table 1, immediately above, reports AdaQualDiff as 30.11 dB / 0.9001 SSIM. These numbers differ by 0.91 dB and 0.009 SSIM, and the claimed ~1.4 dB gain over OneRestore† (28.72) depends on which number is correct. No code, model weights, or evaluation script is released, so a reader cannot determine which figure is the reproducible result. Section 5.2 also claims state-of-the-art on 'specific weather removal tasks,' but Table 2 shows Outdoor-Rain PSNR 31.81 below T3-DiffWeather's 31.99, so the scope of the headline claim is overstated. Since the central contribution is an empirical performance assertion, an unreconciled discrepancy in the primary metric and the lack of a release make the claim unverifiable as stated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"AdaQual-Diff proposes a diffusion-based all-in-one image restoration method in which prompt complexity is modulated by spatially local quality scores. The quality map from DeQAScore is partitioned into regions; for each region, the mean quality is mapped through a linear inverse relation (Eq. 6) to a prompt length, and prompts are selected from high/low pools by feature similarity (Algorithm 1). The training loss adds a quality-weighted noise term and a region-selective perceptual term (Eqs. 7-10). Experiments on CDD-11 and weather benchmarks report 30.11 dB/0.9001 SSIM on CDD-11 (Table 1) and claim state-of-the-art performance on composite degradation and specific weather removal tasks, with 61.12M parameters and two sampling steps. The paper also claims a theoretical connection between perceptual quality and optimal guidance complexity, though the connection is presented through a postulated linear mapping rather than a derived optimality result.","tokens_in":20389,"tokens_out":4528,"duration_ms":43444,"significance":"The central idea of replacing a fixed global prompt with a quality-driven spatially varying prompt, and weighting the loss by local quality, is timely and, if implemented as described, has interesting practical value. The ablation tables (Tables 3-5, supplementary Table 8) are internally consistent and show a clear advantage of the adaptive scheme over fixed-length prompts, and the reported efficiency (two sampling steps, 17 ms per image) is attractive. The paper's contributions are weakened, however, by an unreconciled numerical discrepancy in the headline CDD-11 result, an overbroad state-of-the-art claim on weather removal, and the absence of a derivation that would justify the 'theoretical framework' language. These issues must be resolved before the experimental claims can be fully accepted.","major_comments":[{"comment":"The text states that AdaQual-Diff 'outperforms the previous best method by a large margin with 31.02 dB PSNR and 0.9091 SSIM,' while Table 1 reports 30.11 dB/0.9001 SSIM for AdaQual-Diff and 28.72/0.8821 for OneRestore†. The two versions of the headline result differ by 0.91 dB and 0.009 SSIM, and the claimed margin over OneRestore† changes from roughly 2.3 dB to 1.4 dB depending on which number is correct. Please identify the correct set of numbers and reconcile all occurrences; as written, the central quantitative claim is ambiguous.","section":"§5.2 and Table 1"},{"comment":"The claim that AdaQual-Diff 'achieves state-of-the-art performance on specific weather removal tasks' is not supported by Table 2: on Outdoor-Rain the PSNR is 31.81, below T3-DiffWeather's 31.99, and on RainDrop the SSIM is 0.9330, below T3-DiffWeather's 0.9411. Please qualify the claim to name the datasets and metrics on which the method is actually best, or present the comparison with a more precise scope.","section":"§5.2 and Table 2"},{"comment":"The contribution statement says the paper 'establishes a theoretical framework that formally connects perceptual quality assessment with optimal guidance complexity,' but the paper does not derive Eq. (6) from an optimality criterion; it postulates a linear inverse mapping between quality and prompt complexity. The free parameters C_min, C_max, tau, lambda1, and lambda2 are selected by ablations on CDD-11 validation, and supplementary Table 8 shows a 1.2 dB swing between tau=1.5 and tau=4.5. Please either provide a formal derivation or justification for the inverse linear relation, or rephrase the contribution as a design heuristic, and report how the selected tau transfers to the other benchmarks rather than only to CDD-11.","section":"§1, Eq. (6), Supplementary Table 8"},{"comment":"Algorithm 1 reduces each partition region to its mean quality q_r (step 7) and selects prompt content only through the magnitude C_p and a binary high/low pool decision (step 9), discarding all spatial structure and degradation identity within a region. If two regions with the same mean score but different degradation types require different prompt content, this mechanism cannot deliver the 'precise restoration directives' claimed. Please add an experiment or analysis that tests whether the mean statistic suffices, for example by comparing against per-pixel quality-conditioned prompting or by reporting the frequency with which regions mix different degradation types.","section":"Algorithm 1"},{"comment":"No code, model weights, or evaluation scripts are provided, so the conflicting CDD-11 numbers and the weather benchmark results cannot be independently checked. I recommend releasing these artifacts with the revision, at minimum for the CDD-11 comparison, so that the reported PSNR/SSIM values can be reproduced by readers.","section":"Availability"}],"minor_comments":[{"comment":"The text reads 'across different different adverse weather conditions'; the duplicated word should be removed.","section":"§5.3"},{"comment":"The wording 'Our analysis reveals a mathematical relationship' overstates what is in fact a postulated linear mapping; consider softening this to 'we hypothesize and implement a relationship' until a derivation is supplied.","section":"Fig. 1 caption and §4"},{"comment":"The row labels are ambiguous: the column headers say L1, Prompt, and Quality-weighted, but the checkmarks do not make clear which row corresponds to 'Full Loss (Ours)' other than all boxes being checked; please make the rows self-explanatory.","section":"Table 5"},{"comment":"Eq. (9) uses w(q)=(q_max-q)/(q_max-q_min), which is zero at q_max, and DeQAScore is described as producing values in [1,5]; please clarify whether q_min and q_max in Eq. (6) and Eq. (9) are fixed dataset bounds (1 and 5) or are recomputed per image.","section":"Eq. (9) and §3"},{"comment":"The supplementary text states that the RainDrop-A test subset contains 58 images, while the main paper reports results on 'RainDrop' without specifying the subset; please align the terminology and specify the exact test split used.","section":"Supplementary Section B"}],"recommendation":"major_revision","confidential_remarks":"The mismatch between the prose (31.02/0.9091) and Table 1 (30.11/0.9001) is the kind of issue that, if unresolved, will be noticed by readers and may undermine trust in the other empirical numbers. I lean toward major revision rather than rejection because the adaptive-versus-fixed ablation evidence is coherent and the approach is plausible; however, the authors need to correct the inconsistent numbers, soften the theoretical claim, and ideally release code and checkpoints before acceptance. The arXiv submission also contains placeholder symbols and formatting artifacts, so a careful proofreading pass is needed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know about this paper: the core idea is real and the ablations support it, but the headline number is internally inconsistent and the SOTA claim is overstated.\n\nWhat is new: AdaQual-Diff modulates prompt complexity (token length) regionally based on DeQAScore quality maps, so low-quality regions get long, detailed prompts and high-quality regions get short, preservative ones. That is genuinely different from PromptIR and T3-DiffWeather, which use fixed or degradation-type prompt pools. Algorithm 1 is a clear enough recipe, and the ablations in Tables 3-5 are internally consistent: fixed prompt lengths lose to the adaptive version, DeQAScore beats Compare2Score and Q-Align for guidance, and the quality-weighted loss helps. The supplementary limitations section is honest about computational cost.\n\nNow the soft spots, in order of severity. First, the paper cannot be verified as written. Section 5.2 says AdaQual-Diff outperforms the previous best by a large margin with 31.02 dB / 0.9091 SSIM on CDD-11, but Table 1, a few lines above, reports 30.11 / 0.9001. That 0.91 dB gap changes the claimed 1.4 dB gain over OneRestore† by a third. No code, weights, or evaluation scripts are released, so a reader cannot tell which figure is the real result. That alone should force a revision before acceptance. Second, the \"state-of-the-art on specific weather tasks\" claim is contradicted by the authors' own Table 2: on Outdoor-Rain they get 31.81 PSNR, below T3-DiffWeather's 31.99. Smaller issues: the \"mathematical relationship\" in Eq. 6 is an asserted linear mapping, not a derivation, so the introduction overstates the theory; the threshold τ is tuned on the same dataset family as the headline result (Table 8 shows a 1.2 dB swing around the selected value); and the 17 ms inference time appears to exclude DeQAScore evaluation, which is not free, though the caching discussion acknowledges the issue.\n\nWho this is for: anyone working on all-in-one restoration or quality-guided diffusion conditioning. The adaptive prompting direction is worth thinking about even if this specific implementation needs work. I would send it to peer review, but a serious referee should demand a reconciled number, a release (at least the evaluation scripts), and a scaled-back SOTA claim. As it stands, the central empirical claim is unverifiable.","headline":"Adaptive prompt-length modulation by regional quality scores is a genuinely new and plausible mechanism, but the paper currently can't be verified because its own Table 1 and Section 5.2 quote different numbers for the headline CDD-11 result.","tokens_in":21026,"tokens_out":839,"would_cite":false,"duration_ms":9845,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that in diffusion-based restoration, prompt length should be set by local perceptual quality: the worse a region scores, the longer and more precise its restoration prompt.","keywords":["image restoration","diffusion models","all-in-one restoration","image quality assessment","adaptive prompting","spatially varying degradation","quality-aware conditioning","DeQAScore"],"falsifier":"Build a test image whose regions are matched for mean DeQAScore but differ in degradation type and required operation, such as haze versus rain streaks. If an oracle that knows the true degradation per region produces different prompts and better restorations than AdaQual-Diff's quality-only selection, the sufficiency claim fails; a second check is that the reported swing from 28.91 to 30.11 dB between $\\tau=1.5$ and $\\tau=3.0$ makes the threshold robustness directly measurable on other datasets.","tokens_in":19896,"feed_emoji":"🖼️","tokens_out":5907,"duration_ms":54144,"temperature":0.7,"pith_summary":"AdaQual-Diff claims that the right amount of guidance for a diffusion-based restorer is set by local perceptual quality: areas that score badly should receive longer, more detailed restoration prompts, and areas that already look good should get short, preservation-focused prompts. The paper implements this as an explicit inverse mapping from a DeQAScore quality map to prompt complexity, plus a quality-weighted training loss, inside a two-step diffusion model. If the claim holds, it means a single all-in-one model can adapt to mixed, spatially varying degradations without extra parameters or extra sampling iterations. Reported results include 30.11 dB PSNR and 0.9001 SSIM on CDD-11, ahead of the OneRestore baseline, and competitive scores on snow, rain, and raindrop benchmarks.","feed_headline":"Adaptive prompt lengths lift composite restoration to 30.11 dB","feed_subtitle":"Regions that score worst get longer, more detailed guidance; the mix beats fixed-length all-in-one baselines.","key_machinery":"The load-bearing mechanism is the Adaptive Quality Prompting algorithm (Algorithm 1) built on the inverse-complexity mapping of Eq. 6. It takes a DeQAScore spatial quality map $Q(y) \\in [1,5]^{H\\times W}$, partitions the image adaptively, and for each region reduces the map to its mean score $q_r$; that single scalar sets the prompt token count $C_p$, selects a high- or low-quality prompt pool via threshold $\\tau$, and picks the top-$C_p$ prompts by feature similarity with the region. This selected prompt set then conditions a two-step diffusion NAFNet, and a quality-weighted loss concentrates training gradients on low-quality regions.","core_discovery":"The central discovery is a formal relationship between perceptual quality and guidance complexity in diffusion restoration, stated as prompt complexity scaling inversely with local quality: $C_p(q) = C_{\\min} + (C_{\\max} - C_{\\min}) \\cdot \\left(1 - \\frac{q - q_{\\min}}{q_{\\max} - q_{\\min}}\\right)$. The authors argue that a quality map from DeQAScore, which scores regions on a 1–5 scale, is not just an evaluation metric but an active conditioning signal. Their Adaptive Quality Prompting partitions the image, reads the mean score per region, computes a prompt length from this formula, and selects that many prompt tokens by feature similarity from a quality-appropriate pool, so severely degraded regions get elaborate, targeted prompts while clean regions get minimal preservation prompts. Combined with a quality-weighted loss that upweights errors in low-quality regions, this yields a spatially varying guidance field inside a diffusion model with a NAFNet backbone, and the paper reports that it outperforms fixed-length prompting and prior all-in-one baselines on composite degradations.","pith_inferences":["I infer that the linear inverse mapping is a design choice, not a derived optimum; an experiment comparing Eq. 6 against nonlinear and learned mappings on CDD-11 would show whether the claimed mathematical precision is doing the work or whether any increasing function of $1 - q$ would do.","I infer the same complexity-allocation principle could be applied to computation itself—more diffusion steps, larger receptive fields, or more channels in low-quality regions—which would test whether prompt length is the mechanism or just a proxy for attention.","I infer that collapsing each region to its mean quality score is the fragile step: a region containing both heavy haze and a sharp object gets an averaged prompt that may neither de-haze fully nor preserve detail, so a distribution-aware or degradation-aware regional descriptor is a natural extension.","I infer the caching of quality maps across training iterations assumes the degraded input's quality field is stable; for video or streaming restoration, recomputing spatial quality each frame may be needed, and the cost of DeQAScore would then dominate inference."],"forward_implications":["Fixed-length prompts become a strictly weaker design: the ablation table shows adaptive lengths beat C=10, C=20, and C=30 uniformly, so any all-in-one restorer using constant guidance leaves quality on the table.","Quality assessment stops being a post-hoc evaluator and becomes part of the generative loop; other restoration frameworks can adopt the same score-to-guidance pattern by swapping in their own quality estimator.","Spatial heterogeneity can be handled without extra parameters or extra inference iterations: the same two-step diffusion model allocates effort by prompt selection rather than by running more steps in bad areas.","A model trained on composite degradations transfers to single-weather benchmarks, suggesting the quality-driven prompts do not overfit to one degradation type.","Downstream tasks benefit: object-detection confidence scores on degraded regions improve after restoration, which matters for autonomous-driving and surveillance use cases."],"supporting_citations":[{"why":"Supplies the DeQAScore spatial quality maps that drive the inverse prompt-complexity mapping; without it the method has no quality signal.","marker":"[56]"},{"why":"Provides the CDD-11 composite-degradation dataset and the OneRestore baseline that AdaQual-Diff is designed to beat.","marker":"[15]"},{"why":"Establishes the prompt-based all-in-one restoration paradigm that AdaQual-Diff extends from fixed to quality-adaptive prompts.","marker":"[34]"},{"why":"Recent prompt-pool diffusion restorer (T3-DiffWeather) used as the closest compared baseline and the strongest diffusion competitor.","marker":"[6]"},{"why":"WeatherDiffusion sets the diffusion-based adverse-weather restoration baseline and the global-conditioning approach this paper argues is insufficient.","marker":"[32]"},{"why":"Contributes the TransWeather/AllWeather training data and the all-in-one weather baseline that anchors the benchmark comparisons.","marker":"[42]"},{"why":"Supplies the NAFNet backbone that AdaQual-Diff embeds in its two-step diffusion process.","marker":"[3]"},{"why":"Restormer is the strong regression-based restoration baseline on both composite and weather benchmarks.","marker":"[57]"}],"fun_headline_variants":["Quality scores set prompt length for sharper restorations","Adaptive prompting tailors diffusion to each image region","Prompt detail scales with local degradation in diffusion","No extra parameters: quality-guided prompt lengths win"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The system assumes that a single mean DeQAScore per region is a sufficient statistic for choosing the right restoration prompt, and that one fitted threshold separates the two prompt pools.","fun_headline_variants_meta":{"raw":{"variants":["Quality scores set prompt length for sharper restorations","Adaptive prompting tailors diffusion to each image region","Prompt detail scales with local degradation in diffusion","No extra parameters: quality-guided prompt lengths win"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000394,"raw_usage":{"total_tokens":2079,"prompt_tokens":970,"completion_tokens":1109,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":586,"completion_tokens_details":{"reasoning_tokens":1049}},"tokens_in":586,"tokens_out":1109,"duration_ms":8940,"temperature":1.0,"reasoning_tokens":1049,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:29:31.440405+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a test image whose regions are matched for mean DeQAScore but differ in degradation type and required operation, such as haze versus rain streaks. If an oracle that knows the true degradation per region produces different prompts and better restorations than AdaQual-Diff's quality-only selection, the sufficiency claim fails; a second check is that the reported swing from 28.91 to 30.11 dB between $\\tau=1.5$ and $\\tau=3.0$ makes the threshold robustness directly measurable on other datasets.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the DeQAScore spatial quality maps that drive the inverse prompt-complexity mapping; without it the method has no quality signal."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the CDD-11 composite-degradation dataset and the OneRestore baseline that AdaQual-Diff is designed to beat."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the prompt-based all-in-one restoration paradigm that AdaQual-Diff extends from fixed to quality-adaptive prompts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Recent prompt-pool diffusion restorer (T3-DiffWeather) used as the closest compared baseline and the strongest diffusion competitor."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"WeatherDiffusion sets the diffusion-based adverse-weather restoration baseline and the global-conditioning approach this paper argues is insufficient."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes the TransWeather/AllWeather training data and the all-in-one weather baseline that anchors the benchmark comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the NAFNet backbone that AdaQual-Diff embeds in its two-step diffusion process."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Restormer is the strong regression-based restoration baseline on both composite and weather benchmarks."}],"review_version":1}