{"id":"babe423d-95ed-4ab2-a9da-d7bd9a2b08d4","arxiv_id":"2508.19324","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A survey of deep data hiding methods for ICAO-compliant face images concludes that only a subset of current deep watermarking and steganography models meet the combined requirements of imperceptibility, selective robustness, and blind extraction.","lead":"This survey reviews deep learning watermarking and steganography methods for protecting ICAO-compliant facial images such as passport photos, arguing these techniques can complement capture-time anti-spoofing. It proposes a taxonomy and compares existing methods to identify which can stay invisible, survive compression, still catch tampering, and fit real-world identity systems.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2's cross-method ranking is not externally valid for ICAO face images: metrics come from heterogeneous non-face datasets/payloads, with no JPEG2000 or face-recognition testing, so the §5.2 claim that only a subset is suitable overreaches the evidence.","rationale":"The reader correctly identifies that Table 2 assumes comparability of self-reported metrics across heterogeneous settings; that is a real weakness. My stress-test sharpens it into a broader external-validity problem: the table not only mixes incomparable conditions but evaluates on non-face/non-ICAO data, omits the ICAO-specified JPEG2000 compression, and does not test the actual biometric pipeline or the morphing threat. These are not minor caveats; they are exactly the conditions the central claim is about. The paper itself acknowledges this gap in §6, listing interaction with face recognition and ICAO compliance verification as future work. That limits the survey to a useful taxonomy and a plausible but unverified hypothesis, not an established comparative result. Because the discussion is explicitly qualitative and the authors flag the need for validation, CONDITIONAL remains the right verdict; the load-bearing concern does not demand rejection but should be stated as a limitation before the strong conclusion is accepted.","tokens_in":16071,"tokens_out":4455,"duration_ms":51781,"concrete_test":"Run a controlled benchmark on an ICAO-compliant face dataset (e.g., 500 frontal portraits from ONOT or FaceSigns). For representative methods (SteganoGAN, UDH, MBRS, RIS, StegFormer, FaceSigns), use official weights or retrain at a uniform payload (1 BPP random bitstring), then apply ICAO-style encoding: JPEG2000 at 300 dpi and JPEG QF=90. Measure cover/stego PSNR and SSIM, decode BER, tamper-detection AUROC under a morphing attack, and face-verification TAR at FAR=1e-4 using ArcFace. If the ranked order or the set of passing methods changes, the §5.2 conclusion is not supported by the current table.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—'only a subset of models meet the combined requirements of visual conformity, selective robustness, and reliable decoding' and that INNs 'emerge as the most promising' (§5.2, §6)—rests entirely on Table 2. That table compiles self-reported metrics from papers trained and tested on different domains (ImageNet, COCO, DIV2K, LFW, CelebA), resolutions (128–1024 px), and payloads (0.00065–24 BPP), mixing binary-string hiding with whole-image hiding. PSNR and BER are not commensurable across such varied conditions: a 40 dB PSNR at 24 BPP on natural images is not comparable to 40 dB at 0.0039 BPP on faces. None of the rows evaluates JPEG2000, the codec ICAO specifies for logical storage, nor does any row measure impact on face recognition performance. Only FaceSigns tests semantic manipulation (face swap); no method is assessed under morphing, the central threat motivating the paper. The authors concede the table is qualitative and call for ICAO-specific validation in §6, but the strong comparative conclusion is stated as a finding. Thus the evidence does not yet establish which methods, if any, would actually satisfy ICAO portrait constraints.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper surveys deep learning-based data hiding (watermarking and steganography) for ICAO-compliant face images. It reviews ICAO portrait quality constraints and real-world applications, argues that PAD provides no post-capture protection, and positions fragile/semi-fragile data hiding as a complementary integrity mechanism. The core technical content is a taxonomy of deep architectures (encoder–decoder, GAN, transformer, invertible neural network, diffusion) and a comparative table (Table 2) of fourteen methods across imperceptibility, robustness, capacity, and a subjective overall grade. The central claim, stated in §5.2 and echoed in §6, is that only a subset of current models meet the combined requirements of visual conformity, selective robustness, and reliable decoding, and that INN-based methods are the most promising for ICAO-oriented integrity verification.","tokens_in":16377,"tokens_out":4355,"duration_ms":47176,"significance":"If the comparative conclusion were robust, the paper would provide useful and timely guidance for deploying tamper-evident signals in e-passports, digital travel credentials, and KYC systems. The survey is genuinely novel in targeting ICAO-constrained face images rather than generic media, and it offers a clear problem formulation, a useful taxonomy, and a sensible discussion of why fragile/semi-fragile behavior is needed. The authors also deserve credit for being explicit that the comparison is qualitative and that ICAO-specific validation is still missing. However, the evidence base assembled in Table 2 does not support the strength of the central claim as currently worded. The cross-method comparison mixes self-reported metrics from heterogeneous datasets, resolutions, payloads, and embedding modalities, includes no JPEG2000 or face-recognition evaluation, and assigns grades without a transparent rubric. These issues are fixable by reframing the claims as qualitative tendencies or by adding a controlled ICAO-oriented benchmark, but they currently affect load-bearing conclusions.","major_comments":[{"comment":"The central finding that 'only a subset of models meet the combined requirements...' rests entirely on Table 2, but the metrics in that table are not commensurable. PSNR and BER are taken from original papers using different datasets (DIV2K, COCO, ImageNet, LFW, CelebA), resolutions (128–1024 px), payloads (0.00065–24 BPP), and task types (binary-string hiding vs. whole-image hiding). A 48 dB PSNR at 24 BPP on natural images is not comparable to 36 dB at 0.00065 BPP on faces. Moreover, ICAO logical storage uses JPEG2000 (Table 1), yet no row evaluates JPEG2000, and no row measures the impact of embedding on face recognition performance. The §6 call for future validation acknowledges this, but §5.2 still states the comparative result as a finding. Please either reframe the conclusion as a qualitative tendency with explicit caveats or add a controlled evaluation on ICAO-like face images wi","section":"§5.2 / Table 2"},{"comment":"The overall 'Grade' column (Low+, Medium, Medium-, High+, etc.) is assigned without a stated scoring rule or weighting. The text in §5.2 says it reflects 'a balanced assessment,' but the reader cannot see how imperceptibility, recovery fidelity, robustness behavior, and resolution compliance are combined. Many cells in the table are also missing (—), including JPEG-robustness values for several methods. Without a transparent rubric, the grades are not reproducible and cannot support the strong statement that certain methods are 'better aligned' with ICAO requirements. Please define the grading criteria or remove the grade column and restrict the discussion to the reported values.","section":"Table 2, 'Grade' column"},{"comment":"Section 5 opens by identifying security as one of the four core evaluation dimensions, but Table 2 contains no security metrics. The text mentions that 'most of the suitable methods report performance against generic steganalysis benchmarks... around 55% or less,' but no table rows, references, or per-method values support this statement. Because §3.2 lists security as desirable and §5.2's conclusion discusses combined requirements, omitting security from the comparison weakens the claim of a comprehensive suitability assessment. Either include the available security evaluations or explicitly bound the comparison to imperceptibility, robustness, and capacity.","section":"§5.1 / §5.2"},{"comment":"The paper's central threat model is morphing and semantic manipulation, but no method in Table 2 is evaluated under morphing. The only method with any semantic-manipulation test is FaceSigns (face swap), and that result is not reported in the table. Robustness to JPEG QF=90/50 is not the ICAO-relevant codec condition, since Table 1 specifies JPEG2000 for logical storage. Consequently, the 'selective robustness' component of the central claim is not actually established for the ICAO scenario. The authors should either present this as an open question or add the missing attack evaluations.","section":"§5.2, selective robustness"}],"minor_comments":[{"comment":"The caption of Figure 2 reports PSNR=45.813 and SSIM=0.9861 at 1 bpp for [76], while Table 2 lists Stegaformer [76] at 3 bpp with PSNR values 43.37/47.31. Similarly, Figure 2 reports PSNR=39.562 for [35] at 24 bpp, while Table 2 lists StegFormer [35] at ~24 bpp with PSNR 56.30/55.45 on DIV2K. Please harmonize the reported conditions and values.","section":"Figure 2 / Table 2"},{"comment":"There is an orthographic inconsistency between 'StegFormer' [35] and 'Stegaformer' [76], which are easily confused. Also, the author list repeats 'Jefferson David Rodriguez Chivata' twice; please correct this.","section":"References / author list"},{"comment":"The sentence 'Section 2 revises the key biometric and technical specifications' should read 'reviews' rather than 'revises.'","section":"Section 2, first paragraph"},{"comment":"Table 1 correctly lists JPEG2000 as the logical-storage compression format, but the robustness discussion concentrates exclusively on JPEG compression. The manuscript would benefit from explicitly stating why JPEG2000 was not included in the comparison or from discussing its expected effect on the surveyed methods.","section":"Table 1 / Section 5.2"}],"recommendation":"major_revision","confidential_remarks":"The paper has a reasonable scope and the taxonomy is useful, but the comparative section as written makes a stronger claim than the evidence supports. I believe the authors can address this by reframing the central conclusion as a qualitative, evidence-limited tendency and by making the grading rubric explicit, or by adding a small controlled evaluation on ICAO-like face images with JPEG2000 and face-recognition metrics. The duplicate author entry and naming inconsistencies should be fixed during revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper's focus on ICAO-compliant face images for deep data hiding is genuinely new, and the authors do a solid job of laying out the relevant standards (Table 1) and the taxonomy of methods in Sections 3 and 4. If you work on watermarking or tamper-evident biometrics, this is a useful map of the literature, and the argument that fragile/semi-fragile methods fit the certification setting better than robust ones is sound.\n\nThe problem is Table 2. The central claim—that only a subset of models meet the combined requirements of visual conformity, selective robustness, and reliable decoding, and that INN-based and transformer-based designs are most promising—rests entirely on that table. But the table mixes methods evaluated on ImageNet, COCO, DIV2K, LFW, and CelebA at resolutions from 128 to 1024, with payloads from 0.00065 to 24 bpp. PSNR and BER are not commensurable across those settings. No method is tested under JPEG2000, which is the codec ICAO specifies for logical storage, and none measures impact on face recognition performance. The stress-test note is right about this overreach. To the authors' credit, they say the table is qualitative and call for ICAO-specific validation in Section 6, so the flaw is more overreach than concealment.\n\nThere are also minor internal inconsistencies: the two 'Stegformer' references ([35] and [76]) are easy to confuse, Figure 2 and Table 2 don't line up perfectly, and the 'Grade' column is subjective. The survey also lacks a systematic search protocol, so the method selection may not be representative.\n\nStill, this is a legitimate synthesis, not a puff piece. The authors are transparent about what they did. I'd send it to peer review; a good referee would ask for clearer comparability or a toned-down conclusion. I'd probably cite it if I were writing on biometric watermarking, though I'd treat the comparison table as a map, not a ranking.","headline":"A useful survey with a genuinely new ICAO lens, but Table 2 is too heterogeneous to support the strong conclusion that only a few methods qualify.","tokens_in":16841,"tokens_out":3430,"would_cite":true,"duration_ms":34483,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Only a subset of deep data-hiding models meets ICAO face-image certification needs, the paper argues.","keywords":["data hiding","ICAO face images","digital watermarking","steganography","invertible neural networks","semi-fragile","tamper detection","biometric certification"],"falsifier":"Run the methods graded 'High' and 'Low' on a single ICAO-compliant face dataset under identical conditions (same resolution, JPEG quality, and face-recognition pipeline), measuring PSNR, BER, and recognition accuracy; if the 'High' methods no longer separate from the 'Low' ones, the survey's central distinction collapses.","tokens_in":16000,"feed_emoji":"🛂","tokens_out":2869,"duration_ms":28551,"temperature":0.7,"pith_summary":"ICAO-standard face photos, used in e-passports and identity verification, remain vulnerable to morphing and deepfake manipulation after capture. This survey argues that embedding tamper-evident signals directly into the image is the natural complement to capture-time defenses. It reviews deep learning watermarking and steganography methods, maps them against ICAO constraints, and concludes that only a subset — chiefly invertible-neural-network and transformer-based semi-fragile methods — balance imperceptibility, selective robustness, and reliable blind decoding. If correct, this narrows the practical design space for persistent integrity verification of standardized biometric images.","feed_headline":"Only a few deep data-hiding models fit ICAO face certification","feed_subtitle":"INN-based semi-fragile watermarking leads the pack for tamper-proof passport photos and KYC.","key_machinery":"The organizing constraint is the ICAO portrait-quality specification: inter-eye distance, frontal neutral face, uniform background, JPEG/JPEG2000 formatting. Against that backdrop, the central identity is the fragility-robustness axis: fragile methods degrade under any alteration, semi-fragile methods absorb benign compression but break under semantic tampering, and robust methods are unsuitable because they would let morphing pass. The comparison uses reported PSNR, BER, payload, and JPEG-QF robustness to grade methods; invertible neural networks are highlighted as promising because they make concealment and reveal symmetric, reversible processes and give fine control over this fragility.","core_discovery":"On the paper's own terms, the central claim is that deep learning-based fragile and semi-fragile data hiding, especially methods built on invertible neural networks, 'emerge as the most promising' for ICAO-compliant integrity verification. The survey establishes that robust, general-purpose watermarking is counterproductive for tamper detection because it tolerates exactly the semantic manipulations (face morphing, swapping) that certification must catch. Comparing thirteen methods on PSNR, bit-error rate, payload, and JPEG-compression robustness, it finds that only a subset simultaneously satisfies visual conformity, selective robustness to benign compression, and dependable message recover","pith_inferences":["The comparison conclusion is directly testable: running the shortlisted models on one ICAO-compliant face dataset, with identical resolution, compression, and a single face matcher, would likely reorder some grades.","The same fragile-vs-robust logic should extend to other standardized biometric images, such as fingerprints or iris, where predictable acquisition formats invite manipulation.","The paper's claim implies that even a very successful robust watermarker would fail morph detection; that is a falsifiable design choice, not a theorem.","Because the survey set of methods is not exhaustive, the identified 'subset' may not be the true frontier; a broader architecture search could uncover better semi-fragile designs."],"forward_implications":["Practitioners building e-passport, digital travel credential, or KYC integrity systems should select semi-fragile, blind-extraction methods rather than robust copyright watermarks.","INN-based and transformer-based architectures, such as RIS and StegFormer, are the most promising candidates to test under real ICAO conditions.","The standard metrics (PSNR, BER, JPEG QF=90/50) are necessary but insufficient; security against steganalysis and impact on face-recognition accuracy need dedicated benchmarks.","The paper's qualitative grades give a preliminary selection tool for deployment, not a strict ranking."],"supporting_citations":[{"why":"Defines the ICAO portrait-quality parameters (inter-eye distance, lighting, compression) that ground the survey's constraint analysis.","marker":"[69]"},{"why":"Supplies the unifying survey of deep data hiding and the taxonomy of watermarking versus steganography that structures the comparison.","marker":"[68]"},{"why":"HiDDeN is the seminal end-to-end deep hiding method against which later approaches are positioned.","marker":"[87]"},{"why":"RIS is the INN-based DCT-domain method that anchors the paper's conclusion that invertible architectures are most promising.","marker":"[39]"},{"why":"MBRS provides the JPEG-robustness behavior and high grade that exemplifies the semi-fragile watermarking approach.","marker":"[32]"},{"why":"StegFormer supplies the transformer-based steganography result used to support the advantage of attention-based embedding.","marker":"[35]"},{"why":"FaceSigns demonstrates semi-fragile watermarking against face swap, supporting the selective-robustness requirement.","marker":"[46]"},{"why":"RIIS shows the INN-based robust invertible steganography path that informs the classification of resilience under distortion.","marker":"[73]"}],"fun_headline_variants":["Deep fragile watermarking wins for ICAO face integrity","INN-based watermarking leads tamper-proof passport photos","Most deep watermarking fails ICAO face certification","Only a few deep models keep ICAO face photos tamper-evident"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The paper's grades treat PSNR, BER, and payload numbers reported in each original paper as directly comparable, even though datasets, image resolutions, and payloads differ across methods; if those numbers are not commensurable, the conclusion about which methods are suitable is not robust.","fun_headline_variants_meta":{"raw":{"variants":["Deep fragile watermarking wins for ICAO face integrity","INN-based watermarking leads tamper-proof passport photos","Most deep watermarking fails ICAO face certification","Only a few deep models keep ICAO face photos tamper-evident"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000165,"raw_usage":{"total_tokens":1057,"prompt_tokens":684,"completion_tokens":373,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":428,"completion_tokens_details":{"reasoning_tokens":306}},"tokens_in":428,"tokens_out":373,"duration_ms":4545,"temperature":1.0,"reasoning_tokens":306,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:56:02.258792+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the methods graded 'High' and 'Low' on a single ICAO-compliant face dataset under identical conditions (same resolution, JPEG quality, and face-recognition pipeline), measuring PSNR, BER, and recognition accuracy; if the 'High' methods no longer separate from the 'Low' ones, the survey's central distinction collapses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the ICAO portrait-quality parameters (inter-eye distance, lighting, compression) that ground the survey's constraint analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the unifying survey of deep data hiding and the taxonomy of watermarking versus steganography that structures the comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"HiDDeN is the seminal end-to-end deep hiding method against which later approaches are positioned."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"RIS is the INN-based DCT-domain method that anchors the paper's conclusion that invertible architectures are most promising."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"MBRS provides the JPEG-robustness behavior and high grade that exemplifies the semi-fragile watermarking approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"StegFormer supplies the transformer-based steganography result used to support the advantage of attention-based embedding."},{"cited_title":"FaceSigns: Semi-Fragile Neural Watermarks for Media Authentication and Countering Deepfakes","cited_arxiv_id":"2204.01960","evidence_quote":"FaceSigns demonstrates semi-fragile watermarking against face swap, supporting the selective-robustness requirement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"RIIS shows the INN-based robust invertible steganography path that informs the classification of resilience under distortion."}],"review_version":1}