{"id":"c96df1c9-f469-479d-a1e7-59d88c0a46e4","arxiv_id":"2502.08836","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper is a structured survey that organizes 28 recent papers on deep learning-based single-image reflection removal into three architectural categories, and reviews datasets, metrics, and challenges.","lead":"This paper surveys deep learning methods for removing reflections from single photographs, grouping approaches into single-stage, two-stage, and multi-stage designs. It also lists public datasets, evaluation metrics, and open problems for researchers entering the field.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Comprehensiveness claim is unverifiable: the 28-paper selection from 8 venues is not reproducible, and the stagnation conclusion rests on that incomplete sample.","rationale":"The paper's main contribution is a comprehensive survey and a critical assessment of the field. The strongest claim is that the review is comprehensive and identifies stagnation. The load-bearing condition for that claim is that the sample of 28 papers is representative of the field's output. The methodology does not establish this: the venue list is narrow, the search is not reproducible, and the full paper list is omitted. Other weaknesses (e.g., Table 2 lists CDR as 2021 while the reference is CVPR 2022; no benchmark comparisons) are real but secondary: they affect accuracy and utility without directly invalidating the survey's core thesis. If the selection audit reveals significant omissions, the survey is still useful as an introduction but cannot claim comprehensiveness, and the stagnation narrative would need to be reframed. This is why the verdict should remain conditional pending the audit.","tokens_in":12476,"tokens_out":3621,"duration_ms":35285,"concrete_test":"Reproduce the selection by running the Section 2 query on a fixed set of databases (e.g., DBLP, Scopus, Google Scholar) expanded to include AAAI, ACM MM, BMVC, ICLR, IJCV, IEEE Access, and arXiv for 2017–2025. Compare the resulting candidate set to the reported 28 papers. If more than five additional papers satisfy the stated inclusion criteria, or if any paper cited as state-of-the-art in prior surveys [12] or [10,11] is absent from the list, the comprehensiveness claim is undercut and the stagnation conclusion should be retracted. Also require the authors to publish the full list of 28 papers and their per-paper screening decisions to make the audit possible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2 describes a search restricted to eight venues with a keyword query but no specific database, no inclusion/exclusion criteria, and no full list of the 28 retained papers. Table 1 displays only 18 methods, so readers cannot audit which papers were included or why. The search window (2017–2025) and venue list omit work at BMVC, AAAI, ACM MM, ICLR, IJCV, IEEE Access, and arXiv; the Amanlou et al. systematic review [12] itself appeared in IEEE Access in 2022, demonstrating that relevant work exists outside the chosen venues. Section 6.1 concludes that 'academic research in this field is stagnating' based on this sample, but an incomplete sample makes that conclusion an artifact of the search rather than a property of the field. The paper's own Section 6.3 admits relevant work may have been missed, yet the Abstract still claims a 'comprehensive review' and the paper does not list the included papers. Because the central contribution is comprehensiveness and the survey's critical assessment depends on the sample, the unverifiable selection is the most load-bearing weakness.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of deep-learning-based single-image reflection removal (SIRR). It describes a bibliographic search restricted to eight high-impact venues, organizes the selected literature into single-stage, two-stage, and multi-stage approaches, presents mathematical formation models (linear, blending-scalar, and non-linear), summarizes public datasets and evaluation metrics, and ends with challenges, future directions, and limitations. The authors claim a comprehensive, critically assessed review and identify a three-fold contribution: summarizing recent work, outlining hypotheses/techniques/datasets/metrics, and highlighting challenges and opportunities.","tokens_in":12710,"tokens_out":4547,"duration_ms":45398,"significance":"If the survey were fully reproducible and internally consistent, it would be a useful entry point for researchers entering SIRR. The paper's organization by architectural stage, its compilation of formation hypotheses, and its dataset table are convenient and generally accurate. The explicit discussion of limitations and future directions is also constructive. However, the main value of a survey in this setting is its comprehensiveness and reliability, and both are currently weakened by the unverifiable paper-selection procedure and by specific factual inconsistencies in the dataset section. These issues do not require new experiments to fix, but they do require substantive revision, so the appropriate outcome is major revision rather than rejection.","major_comments":[{"comment":"The paper-selection procedure is not auditable as reported. The authors state that 28 papers remained after their search, but they never list those 28 papers, and Table 1 tabulates only 19 methods, so a reader cannot verify which papers were included or why. Section 6.1 then concludes that \"academic research in this field is stagnating\" based on this unverifiable sample; Section 6.3 also admits that relevant work may have been missed due to the chosen keywords and databases. Because the abstract's \"comprehensive review\" claim and the stagnation conclusion depend on the sample, this is a load-bearing weakness. The authors should provide a complete list of the 28 retained papers (e.g., as an appendix), state the databases searched and the inclusion/exclusion criteria, and either justify the venue restriction or reframe the review as covering a representative sample rather than a comprehensive corpus.","section":"Section 2; Table 1"},{"comment":"The text and Table 2 directly contradict each other on the composition of the SIR2 and CEIL datasets. The text says that \"SIR2 [10] and CEIL [13], are larger and include both synthetic and real-world images,\" while Table 2 labels SIR2 as \"Real\" only and CEIL as \"Syn\" only. This is a factual inconsistency in one of the paper's central reference tables and must be corrected, with the true composition of each dataset verified against the original sources.","section":"Section 5.2; Table 2"},{"comment":"The year for the CDR dataset is inconsistent: Table 2 lists CDR as 2021, but reference [34] is a CVPR 2022 paper. The authors should verify the year and the citation for CDR, and they should audit the other table entries (e.g., the SIR2+ entry, which cites reference [35] that appears to duplicate reference [11]) to ensure that every dataset's year and source match the cited publication.","section":"Table 2; Reference [34]"},{"comment":"The abstract promises a \"critical assessment\" of single-stage and two-stage methods, but Section 4 is largely a descriptive summary of each method, with no explicit comparative strengths/weaknesses analysis. Similarly, the claim in Section 6.1 that the field is \"stagnating\" is stronger than the restricted sample supports, especially since Section 6.3 concedes that relevant work may have been missed. The manuscript should either add a comparative critical discussion (e.g., a strengths/limitations table or a structured comparison of methods) or soften the claims to match the descriptive level actually provided.","section":"Abstract; Section 4; Section 6.1"}],"minor_comments":[{"comment":"There is a typo in the Introduction: \"with the aim to of presenting\" should read \"with the aim of presenting.\"","section":"Section 1"},{"comment":"The text refers to \"DBN\" in the first sentence of Section 4.3, but Table 1 and reference [20] use \"BDN\"; please use the abbreviation consistently.","section":"Section 4.3"},{"comment":"The sentence describing CDR is grammatically broken: \"The CDR dataset is categorized according to reflection types and contains images with perfect alignment between the mixed and transmission images includes misaligned raw flash/ambient images.\" This should be rewritten for clarity.","section":"Section 5.2"},{"comment":"The contextual loss formula in Section 5.3 is presented without an equation number and with an ambiguous \"max_j CX_ij\" expression; please add a numbering and define the index notation.","section":"Section 5.3.2"},{"comment":"References [11] and [35] appear to cite the same TPAMI paper; please merge or disambiguate them to avoid duplication.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a computer-vision venue and addresses a useful topic, but the reproducibility of the literature selection is the key editorial concern: the authors should be required to list the 28 included papers and to clarify the search database and screening criteria. The dataset inconsistencies in Table 2 are also easily fixable but currently undermine confidence in the survey's accuracy. I would not reject the paper, because the taxonomy and discussion have value, but the requested revisions are substantive rather than cosmetic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this survey is a serviceable map of the SIRR deep learning literature, but its 'comprehensive' claim outruns the methodology, and one conclusion (stagnation) leans on a sample you can't audit. The stage-based taxonomy and the dataset summary are the best parts; the selection process is the soft spot.\n\nWhat's new: it updates Amanlou et al. (which stopped at 2021) with papers through 2024, and organizes the field into single-stage, two-stage, and multi-stage architectures plus a discussion of linear vs nonlinear image formation models. That's a real contribution, albeit an organizational one. For a newcomer, Table 1 and Section 5 are genuinely useful.\n\nWhat it does well: the writing is clear, the loss-function section is a nice primer, and the dataset table (CEIL, SIR2, CDR, RRW, etc.) gives a quick sense of scale and real/synthetic balance. The authors also acknowledge in 6.3 that some related work may have been missed and that they did not include benchmark comparisons—both honest disclosures.\n\nWhere it's soft: Section 2 describes a keyword search over eight venues, but gives no database, no inclusion/exclusion criteria, and no list of the 28 papers. Table 1 shows only 18 methods, so the reader cannot verify the selection or the claim of comprehensiveness. That makes the Section 6.1 conclusion that 'academic research in this field is stagnating' an artifact of the sample rather than a demonstrated property of the field. The paper's own limitation statement undercuts the abstract's 'comprehensive review' wording. These are fixable: add a PRISMA-style flow, list all 28 papers, and soften 'comprehensive' to 'focused.'\n\nThere are also small factual slips: the text says SIR2 and CEIL contain both real and synthetic images, but Table 2 labels SIR2 as Real and CEIL as Syn; and CDR's year is 2021 in the table but 2022 in reference [34]. Both should be corrected but neither is load-bearing.\n\nVerdict: this is a decent survey that deserves a serious referee. I'd send it to review with a request to tighten the methodology and fix the inconsistencies, not a desk reject. The reader's conditional verdict is about right; I'd be a bit more charitable on the taxonomy but equally bothered by the reproducibility gap.\n\nWho it's for: grad students or researchers entering SIRR who want a structured overview and a dataset cheat-sheet. I'd cite it as a recent survey reference, and I'd probably bring it to our reading group to discuss what a 'comprehensive' survey owes its readers.","headline":"A useful but uneven survey: the stage-based taxonomy and dataset summary are genuinely handy, yet the 28-paper selection is not auditable and the stagnation claim outruns the evidence.","tokens_in":13189,"tokens_out":2465,"would_cite":true,"duration_ms":23504,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that deep learning research on single-image reflection removal can be organized by image-formation hypothesis and network stage count, and that the field's progress is currently bottlenecked by scarce, low-quality data…","keywords":["reflection removal","single-image reflection separation","deep learning survey","image layer decomposition","synthetic and real datasets","evaluation metrics","computer vision"],"falsifier":"A concrete check: run the same search query across the same eight venues for 2017-2025 and compare the resulting set to the 28 papers the survey says it includes. If any additional relevant paper appears, or if the survey's own references list papers from those venues that are not among the 28, the comprehensiveness claim is falsified.","tokens_in":12299,"feed_emoji":"🪟","tokens_out":6996,"duration_ms":60254,"temperature":0.7,"pith_summary":"This survey tries to establish that deep learning research on single-image reflection removal, published in eight major venues from 2017 to 2025, can be comprehensively organized by two axes: the mathematical hypothesis of how a reflected image is formed, and the number of stages in the network architecture. It argues that the field's progress is currently held back less by network design than by the scarcity of large, diverse, real-world training and test datasets, and by an unclear task definition that mixes reflection removal with background reconstruction. A sympathetic reader would care because the survey provides a structured map of 28 methods, the datasets they use, and the metrics used to judge them, making it easier to position future work and to see where the bottlenecks are.","feed_headline":"Deep learning reflection removal is data-starved, survey finds","feed_subtitle":"A taxonomy of single-, two-, and multi-stage networks reveals that scarce data, not architecture, is the field's bottleneck.","key_machinery":"The organizing device is a two-axis framework. The first axis is the formation hypothesis: how a mixed image $I$ is modeled from a transmission layer $T$ and a reflection layer $R$, ranging from linear $I = T + R$ and $I = \\alpha T + \\beta R$ to non-linear $I = W \\circ T + R$ and $I = T + R + \\Phi(T,R)$. The second axis is the architectural stage count: single-stage networks that map $I$ to $T$ and/or $R$ directly, two-stage networks that first estimate an intermediate (edge map, reflection, or absorption coefficient) and then reconstruct the transmission, and multi-stage recurrent cascades that refine estimates iteratively. This framework is what allows the survey to compare heterogeneous methods and to attribute the field's slowdown to data rather than to model capacity.","core_discovery":"The paper's central claim is that every deep learning approach to single-image reflection removal fits into a taxonomy of single-stage, two-stage, and multi-stage architectures, and that the choice of image-formation hypothesis—from the simple linear superposition $I = T + R$, to blending-scalar variants $I = \\alpha T + \\beta R$, to non-linear $\\alpha$-matte models $I = W \\circ T + R$ and residual formulations $I = T + R + \\Phi(T,R)$—determines how the network is structured and what it can separate. The paper further claims that the current state of the field is characterized by stagnation in network evolution because small datasets make simple models like UNet competitive with elaborate architectures, and that the remedy is larger real-world datasets, clearer task definitions, and the eventual integration of multimodal and foundation-model guidance. This is asserted as a synthesis of the surveyed literature rather than as a new algorithm.","pith_inferences":["The paper does not compare benchmark numbers, so a natural next step would be to run the 28 surveyed methods on a shared test set; this survey provides the roster but not the rankings.","Reading the survey's stagnation argument as a hypothesis, one testable prediction is that a recent model trained on RRW and evaluated on SIR2 will outperform the same model trained on CEIL by a larger margin than the total gains from any architectural change since 2017.","The survey's emphasis on foundation models suggests that language-guided separation, as in the 2024 paper it covers, may become the dominant paradigm; that is an extrapolation the authors hint at but do not develop.","Because the search was limited to English-language venues, the survey's coverage claim implicitly assumes that no significant SIRR work appears in other venues or languages; a reader planning to rely on its completeness should verify against a broader search."],"forward_implications":["If the taxonomy is accepted, new reflection-removal papers can be described by their formation hypothesis and span count, making method comparison more systematic.","Because the survey argues that small datasets let simple UNet-style models reach state-of-the-art, it follows that publishing larger paired real-world datasets would do more for progress than further architectural innovation.","The paper's call for a unified evaluation framework implies that cross-dataset benchmarks on SIR2, CDR, and RRW would settle which methods actually generalize.","The survey's identification of task ambiguity suggests that deciding whether severe reflections require inpainting, not just layer separation, is a prerequisite for meaningful metric comparison."],"supporting_citations":[{"why":"It supplies the first neural network model for SIRR and anchors the survey's 2017 starting point.","marker":"[13]"},{"why":"It introduces the SIR2 benchmark dataset, which the survey's dataset section relies on as a standard test set.","marker":"[10]"},{"why":"It is the previous survey with 2015-2021 coverage, defining the gap that this survey extends.","marker":"[12]"},{"why":"It is the representative single-stage method (ERRNet) around which the survey structures its single-stage discussion.","marker":"[9]"},{"why":"It introduces the residue term that breaks the linear assumption, a key example in the non-linear hypothesis section.","marker":"[16]"},{"why":"It is the canonical multi-stage cascaded refinement approach (IBCLN) that anchors the multi-stage section.","marker":"[21]"},{"why":"It provides the largest real-world dataset (RRW), which supports the survey's data-scarcity argument and future-directions discussion.","marker":"[28]"}],"fun_headline_variants":["Survey: Reflection removal's real bottleneck is data","Reflection removal: data scarcity trumps network design","Deep reflection removal survey flags data shortage","Reflection removal: small datasets stall network evolution","Survey maps reflection removal, pinpoints data gap"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's comprehensiveness rests on the assumption that its keyword search of eight venues, followed by manual screening, captured a representative sample of the field, but the paper does not list the 28 selected papers, so a reader cannot verify that no relevant work was missed.","fun_headline_variants_meta":{"raw":{"variants":["Survey: Reflection removal's real bottleneck is data","Reflection removal: data scarcity trumps network design","Deep reflection removal survey flags data shortage","Reflection removal: small datasets stall network evolution","Survey maps reflection removal, pinpoints data gap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000211,"raw_usage":{"total_tokens":1403,"prompt_tokens":924,"completion_tokens":479,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":409}},"tokens_in":540,"tokens_out":479,"duration_ms":5621,"temperature":1.0,"reasoning_tokens":409,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T23:30:16.597099+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: run the same search query across the same eight venues for 2017-2025 and compare the resulting set to the 28 papers the survey says it includes. If any additional relevant paper appears, or if the survey's own references list papers from those venues that are not among the 28, the comprehensiveness claim is falsified.","supporting_citations":[{"cited_title":"Depth of ﬁeld guid ed reﬂection removal,","cited_arxiv_id":null,"evidence_quote":"It supplies the first neural network model for SIRR and anchors the survey's 2017 starting point."},{"cited_title":"User assisted separation of reﬂec tions from a single image using a sparsity prior,","cited_arxiv_id":null,"evidence_quote":"It introduces the SIR2 benchmark dataset, which the survey's dataset section relies on as a standard test set."},{"cited_title":"Fast single i m- age reﬂection suppression via convex optimization,","cited_arxiv_id":null,"evidence_quote":"It is the previous survey with 2015-2021 coverage, defining the gap that this survey extends."},{"cited_title":"Learning to perceive tr ansparency from the statistics of natural scenes,","cited_arxiv_id":null,"evidence_quote":"It is the representative single-stage method (ERRNet) around which the survey structures its single-stage discussion."},{"cited_title":"Single image reﬂec- tion removal exploiting misaligned training data and netwo rk enhance- ments,","cited_arxiv_id":null,"evidence_quote":"It introduces the residue term that breaks the linear assumption, a key example in the non-linear hypothesis section."},{"cited_title":"Single image reﬂection sep aration with perceptual losses,","cited_arxiv_id":null,"evidence_quote":"It is the canonical multi-stage cascaded refinement approach (IBCLN) that anchors the multi-stage section."},{"cited_title":"Single i mage reﬂection removal through cascaded reﬁnement,","cited_arxiv_id":null,"evidence_quote":"It provides the largest real-world dataset (RRW), which supports the survey's data-scarcity argument and future-directions discussion."}],"review_version":1}