{"id":"85db5b5f-557f-4225-9268-08c315c7b056","arxiv_id":"2506.13307","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A systematic comparison of fine-tuning strategies for Stable Diffusion XL on 100k real SAR images finds that full UNet fine-tuning with LoRA text encoders and a learned <SAR> token gives the best generation quality.","lead":"This paper adapts Stable Diffusion XL to generate radar (SAR) images by comparing full fine-tuning, LoRA adapters, and a learned <SAR> token. A hybrid recipe, full UNet tuning with LoRA text encoders, yielded the best balance of realism and prompt fidelity in their tests.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"KL divergence is used both as a training loss and as the headline evaluation metric, so the reported advantage of the best model is not independent.","rationale":"The reader's CONDITIONAL verdict is appropriate: the paper is a useful empirical study with a plausible winning recipe, but the evidence is not yet conclusive. The reader's weakest_assumption concerns auto-generated optical captions and whether CLIP-based prompt fidelity is meaningful; that is a real risk, and the paper itself acknowledges captioning inaccuracies can cause label noise in Section 3. I regard a different evaluation weakness as more immediately load-bearing: the KL divergence used as the headline statistical-fidelity metric is also part of the training loss for the best model. This is an internally visible circularity, not merely a disagreement with community standards. The concrete ablation test would settle whether the KL advantage is genuine or an artifact of optimizing the evaluation objective. I therefore keep the verdict at CONDITIONAL; no change from the reader's verdict is needed, but the revision requirements should explicitly include removing or independently validating the KL metric overlap and reporting multi-seed variability.","tokens_in":21124,"tokens_out":7820,"duration_ms":83454,"concrete_test":"Ablate the KL term: retrain heart-rose-2 with the identical protocol but λ_KL = 0, so Eq. 10 contains no D_KL(P_real || P_gen) term, and evaluate on the held-out set with KL, CLIP rank, and GLCM. Run the ablation with at least three seeds so that seed noise is bounded. If the KL margin over whale-north-8 largely disappears, the headline 'best' result is an artifact of training on the evaluation metric; if the margin persists, the circularity concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline comparison is partly self-referential. Section 5.5 introduces the refinement loss L_refine = L_base + λ_KL D_KL(P_real || P_gen) (Eqs. 10–11), and Section 5.6 states that the learned <SAR> token is updated using gradients from the diffusion objective and the distributional terms. The best model, heart-rose-2 in Table 6, is therefore trained to minimize the same D_KL that Section 4.3 (Eq. 4) uses as the primary statistical-fidelity metric. Its KL advantage over whale-north-8 (0.23 vs 0.35) may reflect optimization of the evaluation target rather than independent SAR realism, since the loss regularizes exactly the global amplitude histograms that the KL metric compares. The CLIP rank scores provide some independent support, and the GLCM panels are qualitatively useful, but GLCM results are not tabulated per configuration and no confidence intervals are reported for any metric. With one seed and one epoch checkpoint per configuration, the small rank differences among top models (e.g., 1.61 vs 1.74) could be within run-to-run variation. The central claim that the hybrid strategy 'best preserves SAR geometry and texture while maintaining prompt fidelity' is therefore not fully established by the evidence as presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper adapts Stable Diffusion XL to high-resolution X-band SAR imagery. It curates 100k 1024x1024 patches from ONERA SETHI data, auto-generates captions from co-aligned optical images with CogVLM2, and compares fine-tuning configurations that vary full UNet fine-tuning, LoRA on the two text encoders, noise offset, batch size, VAE decoder refinement, and a learnable <SAR> token. Evaluation combines KL divergence between real and generated amplitude histograms, GLCM texture features, and a SAR-specialized CLIP rank/similarity. The central claim is that the hybrid configuration with full UNet tuning, LoRA on both text encoders, and the learned <SAR> token (heart-rose-2) best preserves SAR geometry and texture while maintaining prompt fidelity, and can compose scenes not present in the training set.","tokens_in":21402,"tokens_out":5453,"duration_ms":54990,"significance":"If the result holds, the paper would make a useful contribution: it demonstrates that a text-to-image foundation model can be transferred to a non-optical, physics-driven modality, and it provides a systematic comparison of full fine-tuning versus LoRA across model components. The proposed evaluation framework (KL, GLCM, SAR-CLIP) and the practical applications to TerraSAR-X and EMPRISE refinement are also valuable. However, the headline KL metric is also used as a training loss in the refinement phase, so the most emphasized advantage of the best model is not independent; the single-seed comparisons without uncertainty intervals further limit the strength of the ranking. The paper is within scope and the central idea is defensible, but the evidence as presented does not fully establish the ordering of configurations.","major_comments":[{"comment":"The primary statistical-fidelity metric is used as a training objective. Section 5.5 introduces L_refine = L_base + lambda_KL D_KL(Preal || Pgen), and Section 5.6 states that the <SAR> token embedding is updated with gradients from the diffusion objective and the distributional terms. The evaluation metric in Section 4.3, Eq. (4), is the same KL divergence between real and generated amplitude distributions. The best model heart-rose-2 is therefore trained to minimize the criterion on which it is judged, so its KL advantage over whale-north-8 (0.23 vs 0.35) is at least partially a consequence of optimizing the evaluation target rather than an independent measure of SAR realism. I recommend evaluating at least one configuration trained without the KL term, or reporting KL only as a secondary metric, so that the comparison is not self-referential.","section":"5.5-5.6, Eqs. (10)-(11), Table 6"},{"comment":"All comparisons are single-seed, single-checkpoint runs at epoch 8, with no confidence intervals or significance tests. The differences between top models are small (e.g., CLIP rank 1.61 vs 1.74; KL 0.33 vs 0.35) and could plausibly lie within run-to-run variation. The paper selects the best configuration based on these point estimates, so the central claim is fragile without repeated seeds or bootstrap estimates. Please provide at least 3-5 seeds per configuration, or bootstrap over the 330 evaluation images, and state whether the reported differences are statistically distinguishable. The per-category results in Table 7 likewise lack any uncertainty quantification.","section":"5.2-5.6, Tables 2-6"},{"comment":"The prompt-fidelity evaluation depends on captions whose reliability is not established for the SAR modality. Section 3 states that textual descriptions were automatically generated from the co-aligned optical images using CogVLM2, and the same captions are used as prompts for generation and as ground truth for the CLIP rank and similarity scores in Section 4.3. If the optical captions miss radar-specific content, contain scene-label errors (Section 3 acknowledges possible mislabeling), or use generic phrasing that does not discriminate physical layout, the CLIP scores no longer measure SAR-prompt alignment. Please validate caption quality on a human-annotated SAR subset, or at least quantify caption noise and show that the ranking of configurations is robust to it.","section":"3 and 4.3"},{"comment":"The textural component of the central claim is not quantitatively established across configurations. The GLCM analysis is described in Section 4.3, but the results are presented only as qualitative panels in Figure 15 for heart-rose-2 versus real SAR images, and the caption mentions epoch 10 while the quantitative tables are at epoch 8. There is no per-configuration table of the four Haralick features. Please report the GLCM descriptors with means and standard deviations for all compared models (or at least for the configurations in Tables 2 and 6), and state whether the differences are statistically significant; otherwise the claim that the hybrid strategy best preserves SAR texture is not supported by the evidence as presented.","section":"4.3, Figure 15"}],"minor_comments":[{"comment":"There is a numeric inconsistency for the same configuration: Figure 5 reports rain-beach-6 with KL = 0.17, while Table 2 reports rain-beach-6 with KL = 1.17. Please reconcile these values or clarify that they refer to different checkpoints or evaluation conditions.","section":"5.1 vs 5.2, Figure 5 and Table 2"},{"comment":"The model soleil-up-7 is reported with KL = 0.42 in Table 2 and KL = 0.43 in Table 4 under what appears to be the same configuration and epoch. Please specify the exact evaluation protocol or checkpoint used for each table so that the numbers are reproducible.","section":"Tables 2 and 4"},{"comment":"The caption states '(epoch 10)', while the rest of the experimental section reports results at epoch 8. Please clarify which checkpoint was used for the GLCM analysis and whether the visual comparison is representative of the quantitative comparisons.","section":"Figure 15"},{"comment":"Section 3 says captions were generated only for a subset of geo-aligned optical pairs, while Section 5 refers to a training dataset of 100,000 SAR image-caption pairs. Please clarify how many of the 100k samples have captions and how the remaining samples are used during training.","section":"3 and 5"},{"comment":"The caption of Figure B.2 says 'for 9 different models', but the figure shows only three models (whale-north-8, whale-north-8-refined, heart-rose-2). Please correct the caption.","section":"Appendix B"},{"comment":"The statement 'Contrary to what has been observed in the literature' would benefit from a supporting citation or a more precise description of the specific prior observations being compared.","section":"5.3"},{"comment":"The paper does not state whether code, trained checkpoints, or the evaluation dataset will be released. A data and code availability statement would strengthen the reproducibility claims made in Section 5.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is suitable in scope for the journal and the core experiment is interesting. The main concern is that the headline KL metric is also a training objective, which makes the primary advantage of the best model non-independent; this is fixable by adding a no-KL control or demoting KL in the claims. The absence of uncertainty quantification over seeds and the caption-fidelity issue are also fixable in a revision. I do not see grounds for rejection, but the authors should address the numeric inconsistencies and the statistical robustness before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: if you care about adapting text-to-image diffusion models to non-optical, physics-driven imagery, this paper is worth a look. It compares full fine-tuning, LoRA, and token embedding approaches on a real 100k-patch X-band SAR dataset at 40cm resolution, using SDXL. The winning combination—full UNet fine-tuning, LoRA on the text encoders, and a learned <SAR> token—is believable, and the per-category results are consistent. The dataset and evaluation framework are real assets.\n\nWhat's new: as far as I know, this is the first systematic side-by-side of these fine-tuning recipes for real high-resolution SAR. The individual methods are established, but someone had to run the comparison, and the paper does it carefully in structure: fixed seed, same data ordering, same hyperparameters across runs, per-category evaluation. The noise-offset and LoRA-rank findings are useful practical guidance for anyone trying to repeat this kind of adaptation.\n\nThe soft spots are significant, though. The main one is that the KL divergence used as the evaluation metric (Eq. 4) is also part of the refinement loss (Eqs. 10–11) for the best model, heart-rose-2. So its KL advantage over whale-north-8 (0.23 vs 0.35) is at least partially a consequence of training on that target. That doesn't kill the paper, because the CLIP rank and cosine similarity also favor heart-rose-2, and the qualitative figures look better. But it does mean the 'best preserves SAR geometry and texture' claim rests more on those secondary metrics than the paper acknowledges.\n\nSecond, everything is single-seed, one checkpoint. The rank differences among the top models (1.61 vs 1.74 vs 1.77) are small enough that run-to-run variance could reshuffle them. No confidence intervals anywhere.\n\nThird, the evaluation captions are auto-generated by CogVLM2 from co-aligned optical images. If those captions describe optical properties that don't exist in SAR, the CLIP prompt-fidelity numbers are only measuring alignment to noisy labels. The authors don't validate caption quality for SAR relevance. GLCM texture features are only shown qualitatively for one model, not tabulated per configuration.\n\nAlso, no code or checkpoints are released, which makes the comparison hard to build on.\n\nMy overall take: the central recipe is plausible and the paper deserves peer review, but it needs a revision that either removes the KL term from the training loss or reports a distributional metric that wasn't trained on, adds uncertainty quantification, and validates the captions. The authors are clearly serious about the domain; I'd send it out.","headline":"Plausible hybrid fine-tuning recipe for SDXL-to-SAR, but the headline KL metric is partly trained into the best model and no uncertainty bounds are given.","tokens_in":21931,"tokens_out":3821,"would_cite":true,"duration_ms":36130,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A pretrained text-to-image diffusion model can be tuned to compose realistic, never-before-seen SAR radar scenes.","keywords":["synthetic aperture radar","latent diffusion model","Stable Diffusion XL","LoRA","text-to-image generation","SAR image synthesis","parameter-efficient fine-tuning","SAR-CLIP evaluation"],"falsifier":"Take the 330 evaluation prompts, replace the auto-generated optical captions with independently verified radar-aware captions, and recompute the SAR-CLIP rank and similarity scores for the configurations in Table 2; if the ordering no longer puts the hybrid configurations on top, the reported advantage is an artifact of caption noise. A second check is to verify by geolocation or patch retrieval that the claimed novel scenes, such as a circular structure in a city center or a river through dense forest, are genuinely absent from the training set.","tokens_in":20948,"feed_emoji":"🛰️","tokens_out":5732,"duration_ms":50800,"temperature":0.7,"pith_summary":"This paper tries to establish that a large pretrained text-to-image diffusion model can be moved from ordinary photographs to Synthetic Aperture Radar (SAR) imagery without losing its ability to understand language and compose scenes. The authors adapt Stable Diffusion XL to 100,000 X-band radar patches at 40 cm resolution and systematically compare full fine-tuning against low-rank LoRA adapters in the UNet, text encoders, and VAE. They argue the winning recipe is a hybrid: fully fine-tune the UNet to learn radar-specific speckle, geometry, and amplitude statistics, apply LoRA to the text encoders to keep language knowledge intact, and learn a dedicated <SAR> token embedding. If right, this would make radar imagery a controllable, prompt-driven generative medium, with practical value for data augmentation and rare-scene simulation in Earth observation.","feed_headline":"Hybrid tuning, not full or LoRA, adapts diffusion to radar scenes","feed_subtitle":"Full UNet tuning with low-rank text adapters and a learned token best preserves SAR geometry, texture, and prompt fit.","key_machinery":"The load-bearing object is Stable Diffusion XL's latent diffusion pipeline: a Variational Autoencoder maps images into latents, two text encoders produce token-level and global embeddings, and a UNet denoises latents conditioned on those embeddings through cross-attention. The fine-tuning recipe modifies this pipeline in three modular ways: full gradient updates of the UNet weights, LoRA low-rank adapter matrices on the text-encoder attention projections, and a learnable embedding for a new <SAR> token, plus a short low-noise refinement of the VAE decoder with a KL distribution-matching loss. LoRA updates frozen weights additively with rank-constrained factors, keeping adapter capacity small, while full UNet tuning absorbs the statistical shift that SAR's heavy-tailed amplitude distributions and slant-range geometry impose.","core_discovery":"The central claim is that a vision-language foundation model pretrained on optical images can be adapted to generate physically grounded SAR scenes, and that the optimal configuration is neither all-full nor all-LoRA but a hybrid. Full fine-tuning of the UNet captures SAR's low-level statistics (Rayleigh-distributed amplitudes, speckle, side-looking geometry), while LoRA on both text encoders preserves spatial relations and object semantics, and a learned <SAR> token shifts the caption embedding toward radar style. The paper reports that the hybrid model heart-rose-2 reaches a KL divergence of 0.23 against real amplitude distributions and attains the best SAR-CLIP ranks in most evaluated categories, while composing scenes the training set did not contain, such as a circular structure in a city center or a river through dense forest.","pith_inferences":["Inference: the same division of labor seen here—full backbone tuning for low-level statistics, low-rank adapters for language semantics, and a learned token for domain identity—likely transfers to other non-optical modalities with heavy-tailed statistics, such as sonar, lidar echo, or medical ultrasound.","Inference: because semantic evaluation depends on auto-generated optical captions, re-running the rankings with manually verified or radar-native captions is a natural stress test; if configuration ordering changes, part of the hybrid's reported advantage is an artifact of caption noise.","Inference: the learned <SAR> token probably behaves like textual inversion, meaning 'a <SAR> image of X' should inherit radar statistics even when X is an unseen concept; that compositional property is directly testable.","Inference: the claim that generated scenes are 'unseen' relies on absence from the training split, so checking by geolocation or patch retrieval whether the composed scenes actually appear in training data would sharpen or weaken the out-of-distribution claim."],"forward_implications":["Text prompts alone can control SAR scene composition, including spatial relations like 'near' and 'along', at 40 cm resolution.","The same adapted model can be reused as an image conditioner: it sharpens 1.35 m TerraSAR-X images to 40 cm and adds physically plausible texture to outputs of the EMPRISE physics-based simulator.","Rare or operationally relevant scenes absent from the training set, such as a bridge over a river or a circular structure in a city, can be synthesized on demand.","The adaptation runs on a single H100 GPU, so the procedure is reproducible for other resolutions or other latent diffusion backbones."],"supporting_citations":[{"why":"Supplies the latent diffusion formulation that the entire adaptation pipeline builds on.","marker":"Rombach et al. (2022)"},{"why":"Defines Stable Diffusion XL, the concrete architecture whose VAE, dual text encoders, and UNet are being fine-tuned.","marker":"Podell et al. (2023)"},{"why":"Introduces LoRA, the parameter-efficient low-rank adaptation method compared against full fine-tuning.","marker":"Hu et al. (2021)"},{"why":"Provides the textual-inversion concept behind the learned <SAR> token embedding.","marker":"Gal et al. (2022)"},{"why":"CogVLM2 is used to automatically generate the captions that form the text side of the training and evaluation pairs.","marker":"Hong et al. (2024)"},{"why":"Defines the GLCM texture features used to measure textural realism of generated SAR images.","marker":"Haralick et al. (1973)"},{"why":"The Segment Anything Model provides the segmentation masks that isolate homogeneous patches for texture evaluation.","marker":"Kirillov et al. (2023)"},{"why":"Documents the SETHI radar sensor and platform used to collect the 100,000-image X-band dataset.","marker":"Baqué et al. (2019)"}],"fun_headline_variants":["Full UNet, LoRA text, learned token: best SAR diffusion","Hybrid tuning for SAR: full UNet, LoRA text adapters","Neither full nor LoRA alone: hybrid adapts diffusion to radar","For SAR, mix full UNet tuning with LoRA on text encoders","Optimal SAR diffusion: full UNet, LoRA text, token embedding"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results stand or fall on the assumption that the captions automatically generated from optical imagery are faithful descriptions of the co-aligned SAR patch; if those captions miss radar-specific content, the semantic alignment scores no longer measure how well the generated SAR image matches the prompt.","fun_headline_variants_meta":{"raw":{"variants":["Full UNet, LoRA text, learned token: best SAR diffusion","Hybrid tuning for SAR: full UNet, LoRA text adapters","Neither full nor LoRA alone: hybrid adapts diffusion to radar","For SAR, mix full UNet tuning with LoRA on text encoders","Optimal SAR diffusion: full UNet, LoRA text, token embedding"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000336,"raw_usage":{"total_tokens":1874,"prompt_tokens":969,"completion_tokens":905,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":585,"completion_tokens_details":{"reasoning_tokens":807}},"tokens_in":585,"tokens_out":905,"duration_ms":8642,"temperature":1.0,"reasoning_tokens":807,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:05:13.815869+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the 330 evaluation prompts, replace the auto-generated optical captions with independently verified radar-aware captions, and recompute the SAR-CLIP rank and similarity scores for the configurations in Table 2; if the ordering no longer puts the hybrid configurations on top, the reported advantage is an artifact of caption noise. A second check is to verify by geolocation or patch retrieval that the claimed novel scenes, such as a circular structure in a city center or a river through dense forest, are genuinely absent from the training set.","supporting_citations":[],"review_version":2}