{"id":"c3183993-763d-464c-b078-8ffb91e6660a","arxiv_id":"2501.10325","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A diffusion model that generates latent high-frequency maps from low-quality stereo images and injects them into a transformer restoration network yields modest gains on stereo super-resolution, deblurring, and low-light enhancement.","lead":"DiffStereo uses a diffusion model to generate high-frequency feature maps that guide a transformer network for stereo image restoration. It reports small accuracy gains and some perceptual improvements on super-resolution, deblurring, and low-light enhancement benchmarks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 4's large LHFR gains come from oracle (ground-truth) representations, while Table 5 shows the diffusion-estimated LHFR add only ~0.07 dB on SR; the paper's central mechanism is unconfirmed.","rationale":"I read the paper as claiming that a latent diffusion model estimating high-frequency representations of HQ stereo pairs, fused into a transformer, yields SOTA accuracy and perceptual quality. For that claim to hold, the DM must predict useful LHFR from LQ inputs at inference. The paper's own ablations are the best evidence on this. Table 4 row 3 vs row 1 shows a 3.45 dB jump, but that row uses Stage 1 (oracle LHFR from GT). Table 5, which uses the actual DM-estimated LHFR, shows only 0.07 dB on Flickr SR and 0.06 dB on low-light. This internal inconsistency is not resolved anywhere in the manuscript. The section text explicitly attributes the Table 4 gain to 'latent representations estimated by DM,' which conflates two different pipelines. This is the weakest load-bearing point. A controlled experiment with oracle vs estimated vs no LHFR under identical conditions would settle whether the diffusion component matters. I do not see a need to change the reader's CONDITIONAL verdict; the concern is exactly the one the reader identified.","tokens_in":13110,"tokens_out":3246,"duration_ms":32015,"concrete_test":"Run a controlled comparison on Flickr1024 ×4 SR with identical training budget and backbone: (A) full DiffStereo with DM-estimated LHFR; (B) same model with oracle LHFR extracted from GT at inference; (C) same SIRN without LHFR; (D) a stereo-adapted DiffIR-style baseline where a DM estimates a 1D condition vector instead of LHFR. Report PSNR/SSIM/FID. If (A)−(C) remains ≤0.1 dB while (B)−(C) is >3 dB, the diffusion guidance is not the source of the claimed gains, and the central claim should be reframed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3.1 presents Table 4 (rows 1–4) as evidence that LHFR 'estimated by DM' improve restoration significantly. But all rows list Stage 1, and in Stage 1 the LHFR are extracted from ground-truth HQ images by LREN; the diffusion model is not used. Row 3 vs row 1 gives a 3.45 dB gain on Flickr1024. Table 5, which evaluates the actual inference-time setup with DM-estimated LHFR, shows SR gains of only 0.07, 0.02, 0.05, and 0.11 dB on the four datasets, and 0.06 dB on low-light enhancement. The text in §4.3.1 that 'the latent representations estimated by DM benefit the quality of reconstruction' is not supported by Table 4; the large gain is an oracle-guidance result. Thus the load-bearing assumption that DM-estimated LHFR supply useful high-frequency information to SIRN is currently unverified, and the measured effect is near noise level on the main benchmark. Final SOTA margins are also small on SR (e.g., 0.05 dB over SSRDE-FNet on Flickr1024), so the claimed 'higher reconstruction accuracy' may be driven by the CIB/SIRN backbone rather than by the diffusion model.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes DiffStereo, a two-stage framework for stereo image restoration that combines a diffusion model with a transformer-based restoration network. In Stage 1, a latent representation extraction network (LREN) is trained jointly with a stereo image restoration network (SIRN) to compress high-frequency details of HQ stereo images into single-channel latent high-frequency representations (LHFR), which are fused into the SIRN via a depth-position encoding scheme. In Stage 2, a diffusion model is trained to estimate these LHFR from degraded stereo images and the estimates are used as guidance at inference. Experiments on ×4 super-resolution, deblurring, and low-light enhancement report favorable PSNR/SSIM, FID, and LPIPS against CNN/transformer baselines. The central claim is that this combination yields both higher reconstruction accuracy and better perceptual quality than state-of-the-art methods.","tokens_in":13425,"tokens_out":5730,"duration_ms":49396,"significance":"If substantiated, the work would be the first diffusion-based stereo image restoration method and introduces a plausible design choice: running the diffusion process in a spatially preserved, channel-compressed latent space rather than in pixel space, and using the generated representation as auxiliary guidance rather than as the direct output. This is a sensible way to avoid the typical PSNR drop of generative restoration models. The experimental coverage of three tasks and the inclusion of perceptual metrics (FID/LPIPS) are strengths. However, the paper's own ablations do not currently confirm that the diffusion-estimated LHFR are responsible for the reported gains: Table 4 evaluates oracle (ground-truth) LHFR at Stage 1, while Table 5 shows only marginal DM-based gains. The 'high-frequency' identity of the learned representation is also asserted largely from visualizations rather than from quantitative analysis. The central mechanism is therefore plausible but unverified, and the SOTA margins on SR are small.","major_comments":[{"comment":"The text states that 'the latent representations estimated by DM benefit the quality of reconstruction' and cites Table 4, but all rows of Table 4 are Stage 1 results, in which the LHFR are extracted from ground-truth HQ images by the pretrained LREN and the diffusion model is not used. Row 3 versus row 1 shows a 3.45 dB gain on Flickr1024, which is an oracle-guidance result, not a DM result. The actual inference-time contribution of the diffusion-estimated LHFR is shown in Table 5: on SR the gains are 0.07, 0.02, 0.05, and 0.11 dB on Flickr1024, KITTI2012, KITTI2015, and Middlebury, respectively, and 0.06 dB on low-light enhancement; several SSIM values decrease or stay flat (e.g., deblur Flickr1024 SSIM 0.9790 vs 0.9790, KITTI2012 SSIM 0.9725 vs 0.9724). These magnitudes are close to noise level and do not support the claim that the DM-estimated LHFR are the mechanism behind the reported improvements. The paper should present a Stage-2 ablation with and without DM-estimated LHFR as the main evidence, and clearly separate oracle and estimated settings in Table 4.","section":"§4.3.1, Tables 4 and 5"},{"comment":"The diffusion-related loss is not derived from the standard DDPM objective in Eq. (14). Instead, the network is trained by backpropagating the L1 distance between the final deterministic output after T iterations and the LREN-extracted LHFR, with the random term dropped in Eq. (11). This is a legitimate objective for a deterministic iterative denoiser, but it is not the variational bound of a diffusion model, and the paper does not explain why the full-chain loss is preferable to the conventional noise-prediction loss. The choice of T=4 is justified only by a sentence in the Fig. 6 caption ('more iterations do not provide further improvement') without a supporting ablation. Either derive the objective from a principled surrogate, or reframe the module as an iterative refinement network and provide the missing T-ablation.","section":"§3.3 and Eq. (15)"},{"comment":"The claim that the learned latent representations are 'high-frequency' is not quantitatively established. The LHFR are defined as the output of LREN, which is co-trained with the reconstruction loss of the final restoration; the visualization in Fig. 6 shows edge-like patterns, but no spectral or fidelity analysis is provided to show that these representations capture high-frequency content of the HQ image beyond what a generic compressed feature would capture. Since the method is named and motivated by high-frequency awareness, the paper should include a quantitative characterization (e.g., comparing the Fourier spectrum of the LHFR with that of the input LQ/HQ images, or a detail-recovery metric) and, ideally, an ablation with a similarly sized latent that is not trained with the 'high-frequency' objective.","section":"§3.2.1 and Fig. 6"}],"minor_comments":[{"comment":"The 'Stage' column should be renamed or annotated so that rows 3 and 4 clearly indicate that LHFR are extracted from the ground-truth HQ images (oracle) and not estimated by the diffusion model; the current layout invites the misinterpretation that the text makes.","section":"§4.3.1, Table 4"},{"comment":"The notation for the extracted ground-truth LHFR and the estimated LHFR is confusing because both are denoted Z in different equations; use e.g. Z^gt and Z^est consistently.","section":"§3.4, Eq. (15)"},{"comment":"'interations' should be 'iterations', and the claim about T=4 should be moved to the main text with an ablation table.","section":"Fig. 6 caption"},{"comment":"The column header 'Left (Left + Right) / 2' is ambiguous; it should be split into two settings, e.g., 'Left' and '(Left+Right)/2'.","section":"§4.2.2, Table 2"},{"comment":"There are several typos and unresolved references, including 'denosing' in §3.3, 'generalizaiton' in §4.2.2, 'Visiual' in Fig. 4 caption, and a placeholder '[31?]' in §2.2.","section":"Various"},{"comment":"What is called 'position encoding' is a depth-index concatenation learned via convolution; the name might be confused with spatial position encoding used in transformers. Consider renaming to 'depth conditioning' or justify the terminology.","section":"§3.2.3"},{"comment":"The abstract's claim of 'higher reconstruction accuracy' is not universally supported: on KITTI2012 DiffStereo's PSNR (26.67) is lower than SSRDE-FNet (26.69). The text should state the per-dataset comparison precisely, e.g., as 'comparable on KITTI2012/KITTI2015 and better on Flickr1024/Middlebury'.","section":"Abstract and Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable candidate for a stereo-restoration venue, but the gap between the presentation and the evidence is significant. The authors should be asked to provide a Stage-2 ablation with and without DM-estimated LHFR and to correct the misleading use of Table 4. The 'first diffusion model for stereo restoration' claim is fine, but the actual contribution of the diffusion module needs to be demonstrated. If the authors cannot provide such an ablation, the paper's central mechanism remains unverified and rejection might be warranted; however, within a revision cycle the issue is fixable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Qing,\n\nBottom line: this is the first diffusion-model-based stereo image restoration paper, and the high-frequency spatial latent is a genuine departure from DiffIR's compact vector. But the single most important evidence for the mechanism has a hole: the big gains in the LHFR ablation come from ground-truth maps, not from the diffusion model. In Table 4, rows 1–4 are all Stage 1, where the LHFR are extracted from HQ images by the pretrained LREN. Row 3 vs row 1 is a 3.45 dB jump on Flickr1024—that is oracle guidance. Table 5, which measures the real inference-time setup with DM-estimated LHFR, shows only 0.07 dB on Flickr1024 SR and 0.06 dB on low-light. The text in §4.3.1 claims the DM-estimated latent representations 'benefit the quality of reconstruction,' but the table does not show that. The measured effect is close to noise level, so the load-bearing assumption that the DM supplies useful high-frequency information is unverified.\n\nWhat is genuinely good: the idea of a full-resolution, single-channel latent that keeps spatial detail; the position-encoding scheme across network depth; and the three-task evaluation with FID and LPIPS. The architecture is described well enough to build on. The related work is honest about DiffIR as the closest ancestor, though a direct comparison against a stereo-adapted DiffIR is missing.\n\nThe other soft spots are secondary. Eq. (15) deviates from standard DDPM training—run all T iterations at once and minimize L1 on the final estimate—without derivation. That can be a legitimate training objective, but with T=4 it deserves a word. And the final SOTA margins on SR are tiny (0.05 dB on Flickr1024), so even the headline numbers suggest the diffusion part is not the main driver; the CIB/SIRN backbone may be. The deblurring and low-light gains are larger, but Table 5 shows the LHFR contribution is still small there.\n\nThis paper deserves a serious referee. The application is new, the method is clearly presented, and the issues are fixable if the authors report the oracle-vs-estimated gap explicitly, correct the text about Table 4, and add the DiffIR baseline. I'd send it out with a request for major revision.","headline":"First stereo diffusion restoration paper, but the key ablation confuses oracle guidance with what the diffusion model actually delivers.","tokens_in":13915,"tokens_out":3031,"would_cite":false,"duration_ms":29423,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DiffStereo is the first diffusion model for stereo image restoration.","keywords":["DiffStereo","stereo image restoration","diffusion model","latent high-frequency representation","transformer","stereo super-resolution","deblurring","low-light enhancement"],"falsifier":"A decisive check is to ablate the diffusion-estimated LHFR on the same benchmark by replacing them with a constant map after training; if PSNR/SSIM stay within the roughly 0.07 dB margin seen in Table 5, the high-frequency guidance is not load-bearing, while a large drop would confirm the mechanism.","tokens_in":12914,"feed_emoji":"🖼️","tokens_out":8962,"duration_ms":76948,"temperature":0.7,"pith_summary":"DiffStereo sets out to bring diffusion models to stereo image restoration, a setting where they have not been used, by running the diffusion process on a compact latent space dedicated to high-frequency detail rather than on full images or semantic latents. The model first distills ground-truth stereo pairs into a single-channel latent high-frequency representation (LHFR) at full resolution, then trains a diffusion model to predict that map from the degraded left-right input, and finally fuses the predicted map into a transformer-based restoration network with a depth-dependent position encoding. The paper reports that this design improves both pixel accuracy (PSNR/SSIM) and perceptual quality (FID/LPIPS) over state-of-the-art CNN and transformer methods on stereo super-resolution, deblurring, and low-light enhancement. If the claim holds, DiffStereo is the first effective diffusion-based approach to stereo restoration, and it offers a division of labour where the diffusion model supplies realistic texture guidance while the regression network keeps fidelity.","feed_headline":"DiffStereo: first diffusion model for stereo image restoration","feed_subtitle":"A latent high-frequency map from diffusion guides a transformer to restore stereo pairs sharply.","key_machinery":"The central object is the latent high-frequency representation (LHFR): a single-channel map at the same resolution as the input stereo images, produced by LREN from the HQ pair and estimated by the diffusion model from the LQ pair. It carries the high-frequency texture and edge information that ordinary latent diffusion discards during compression, and its full spatial resolution preserves the texture of the input. The other load-bearing pieces are the channel interaction block (CIB), a transformer unit that computes a shared channel attention map over concatenated left-right features, and the position encoding scheme that modulates CIB features with the LHFR differently according to the block's depth. Together they let the diffusion prior steer fine detail while the transformer enforces cross-view consistency and fidelity.","core_discovery":"The central claim is that the right role for a diffusion model in stereo restoration is not to generate the output image but to generate a compressed, full-resolution high-frequency map of the target stereo pair. The latent representation extraction network (LREN) is trained, jointly with the stereo image restoration network (SIRN), to distill ground-truth stereo pairs into a single-channel map that retains edge and texture structure. A diffusion model is then trained in that space to predict the map from the degraded stereo input, and the predicted map is fused into each channel interaction block (CIB) of SIRN with a position-encoding scheme that lets the same map guide different network depths differently. The paper argues this keeps the generative power of diffusion for distribution-level realism while the regression-based transformer keeps pixel fidelity, avoiding the accuracy drop typical of image-space diffusion restoration. Experiments across three tasks are presented as evidence that this division of labour improves both PSNR/SSIM and FID/LPIPS over previous methods.","pith_inferences":["A cautious inference from the ablations is that the diffusion-estimated maps are not yet the main source of the accuracy gain: on the super-resolution benchmark they add about 0.07 dB, while the oracle-map comparisons show gains of several dB, so the transformer backbone and training-time information are likely doing much of the work.","A testable extension beyond the paper is to condition the diffusion model on disparity or epipolar geometry from the stereo pair, in addition to the degraded views; if high-frequency prediction is the bottleneck, cross-view geometric cues should close part of the gap.","The same LHFR-as-guidance design could transfer to other multi-view settings such as light-field restoration or multi-frame super-resolution, where one shared full-resolution high-frequency map could be estimated by a single diffusion pass and then guide a regression network."],"forward_implications":["Diffusion models can be applied to stereo restoration without the usual compute blow-up, because the diffusion process runs on a one-channel map at full resolution instead of on two full-color images.","Accuracy and perceptual quality do not have to be traded off: on the reported tasks DiffStereo improves PSNR/SSIM and FID/LPIPS at the same time, unlike image-space diffusion methods that typically sacrifice one.","Because the LHFR is estimated from the degraded input and then fused into a regression network, the diffusion model does not directly synthesize pixels, which the paper credits with suppressing hallucinated details.","The position-encoding fusion means the same high-frequency guidance can be reused at different network depths, giving the restoration network a way to exploit the prior at multiple scales of abstraction.","The framework transfers across degradation types: super-resolution, deblurring, and low-light enhancement are all handled by the same architecture, suggesting the LHFR prior is not task-specific."],"supporting_citations":[{"why":"Supplies the denoising diffusion probabilistic formulation that DiffStereo adapts to estimate LHFR.","marker":"[13]"},{"why":"Foundational formulation of diffusion probabilistic models as nonequilibrium thermodynamics, which the paper builds on.","marker":"[37]"},{"why":"Introduces latent diffusion compression, the paradigm that DiffStereo modifies to preserve high-frequency details.","marker":"[30]"},{"why":"Prior work using a diffusion model to generate a compact representation for restoration, which DiffStereo extends to full-resolution spatial maps.","marker":"[48]"},{"why":"Provides the multi-Dconv head transposed attention and gated-Dconv feed-forward network that the channel interaction block is based on.","marker":"[53]"},{"why":"Source of the parallax-related loss functions and a stereo super-resolution baseline that DiffStereo compares against.","marker":"[47]"},{"why":"Introduces the parallax attention mechanism and serves as a core stereo restoration baseline.","marker":"[45]"},{"why":"Supplies the Flickr1024 dataset used for training and evaluation of stereo super-resolution.","marker":"[46]"}],"fun_headline_variants":["First diffusion model for stereo restore, via high-freq maps","DiffStereo: diffusion predicts latent high-frequency details","Stereo restoration: diffusion generates texture maps, not pixels","High-frequency aware diffusion for stereo image restoration","DiffStereo: first DM for stereo, using latent HF guidance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pipeline rests on the assumption that a one-channel map distilled from the ground-truth stereo pair retains the specific high-frequency detail that matters for restoration, and that the diffusion model can predict that map from the degraded input; if the predicted map adds only about 0.07 dB on the main super-resolution table, the claimed mechanism is not what carries the results.","fun_headline_variants_meta":{"raw":{"variants":["First diffusion model for stereo restore, via high-freq maps","DiffStereo: diffusion predicts latent high-frequency details","Stereo restoration: diffusion generates texture maps, not pixels","High-frequency aware diffusion for stereo image restoration","DiffStereo: first DM for stereo, using latent HF guidance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000163,"raw_usage":{"total_tokens":1265,"prompt_tokens":987,"completion_tokens":278,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":603,"completion_tokens_details":{"reasoning_tokens":198}},"tokens_in":603,"tokens_out":278,"duration_ms":3555,"temperature":1.0,"reasoning_tokens":198,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:12:39.776091+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check is to ablate the diffusion-estimated LHFR on the same benchmark by replacing them with a constant map after training; if PSNR/SSIM stay within the roughly 0.07 dB margin seen in Table 5, the high-frequency guidance is not load-bearing, while a large drop would confirm the mechanism.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the denoising diffusion probabilistic formulation that DiffStereo adapts to estimate LHFR."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the multi-Dconv head transposed attention and gated-Dconv feed-forward network that the channel interaction block is based on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the parallax-related loss functions and a stereo super-resolution baseline that DiffStereo compares against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the parallax attention mechanism and serves as a core stereo restoration baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Flickr1024 dataset used for training and evaluation of stereo super-resolution."}],"review_version":1}