{"id":"5db720e0-d65f-4853-bf35-cabe946f5478","arxiv_id":"2607.26729","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"CASIAL couples cover-aware global message spreading with spatial-attention alignment to keep deep image watermarks decodable under severe geometric attacks while improving visual quality.","lead":"CASIAL is a deep image watermarking system that stays readable after heavy cropping, rotation, shear, and jigsaw scrambling by spreading each bit across cover-aware features and realigning them with spatial attention. It matters because geometric attacks still break most learned watermarks used for ownership and provenance.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The geometric-robustness claim rests on a matched white-box noise menu that may overstate gains versus baselines not designed or equally tuned for those exact attacks.","rationale":"The reader correctly flags that white-box geometric noise layers plus attention/VAE latents are asked to carry the geometry-invariance and transfer story without theoretical invariance, sync marks, or adaptive attack evaluation. That is the right soft spot: the paper’s evidence is broad and the CAS/IAL diagnostics (Figs. 4–5, Table 4) are coherent, but the strongest claim is a comparative robustness ranking under a shared, author-chosen attack menu. My concern sharpens the same assumption into a concrete fairness risk on the Table 2 headline rather than a generic “no code / no steganalysis” complaint. Nothing in the text internally falsifies the reported numbers, so REJECT is unwarranted; CONDITIONAL remains appropriate pending fairer re-training checks, variance, and code. Agreement with the reader is therefore agree on the weakest assumption and on the verdict band.","tokens_in":16756,"tokens_out":720,"duration_ms":17628,"concrete_test":"Re-train the two strongest geometric-capable baselines (e.g., VideoSeal, ChunkySeal, and one spreading variant) with an expanded, architecture-matched geometric curriculum (multi-scale crop ratios, continuous rotate/shear, elastic, jigsaw) under the same compute budget and message length; re-evaluate Table 2 geometric columns and average. If their bit accuracies rise within ~3–5 points of CASIAL or the ranking flips on C&R/jigsaw while PSNR stays comparable, the “substantially stronger than eleven baselines” claim weakens to a same-pipeline result.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (abstract; Table 2) is that CASIAL is substantially more robust than eleven retrained baselines under six geometric transforms (erase 80%, C&R retain 20%, shear/rotate averages, elastic α=2, 8×8 jigsaw) while keeping best fidelity (Table 1). That comparison is load-bearing on the Experimental Settings statement that “all models are trained under the same noise settings.” The noise layer is the fixed Kornia white-box menu used both to train CASIAL and to define the test distribution (Motivation & Methods; Noise layer; Supplementary Experimental Settings). Prior END/invertible methods were originally built around signal/JPEG-style layers; forcing the same geometric menu does not guarantee equal capacity, schedule, or hyperparameter fairness, and several baselines collapse near chance on C&R/shear/rotate/jigsaw while CASIAL stays ≥98.44%. Ablations (Table 4) show CAS+IAL are necessary inside this pipeline, but they do not show that the headline margin would survive if baselines received architecture-appropriate geometric augmentation, multi-scale training, or sync-style auxiliaries. Black-box transfer (Table 3) is only to appearance filters (Crayon/Film/Sketch), not to unseen geometric operators or adaptive/steganalytic attackers. Thus the “geometry-invariant representation” claim is still an empirical fit to the training attack family rather than a demonstrated general geometric property.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes CASIAL, an encoder–noise–decoder image watermarking framework aimed at geometric robustness. It attributes geometric failures to region removal and desynchronization, and addresses them with two components: CAS, which generates cover-conditioned candidate features and selects them via message bits to spread watermark evidence globally, and IAL, which uses spatial attention to aggregate and realign features under spatial warps. Encoding/decoding operate in VAE latent space with JND residual attenuation for imperceptibility. Empirically, against eleven retrained baselines under matched resolution, message length, and white-box noise settings, CASIAL reports the best fidelity (Table 1: 40.82 dB PSNR) and near-ceiling bit accuracy on six geometric transforms plus competitive signal/photometric results (Table 2: 99.40% average), with ablations (Table 4), residual/attention visualizations (Figs. 4–5), multi-shape tests (Table S1), and transfer to five appearance-style black-box filters (Table 3).","tokens_in":17118,"tokens_out":1487,"duration_ms":46924,"significance":"Geometric robustness remains a genuine bottleneck for deep watermarking, and the paper’s two-failure-mode framing is clear and useful. The empirical package is comparatively strong for this area: matched retraining of eleven baselines, component ablations, content-coupling visualizations, decoder attention under jigsaw, distortion-strength curves, and multi-aspect-ratio evaluation. If the gains hold under fairer geometric augmentation of baselines and under geometric operators outside the training menu, CASIAL would be a meaningful practical advance for copyright/provenance pipelines that must survive crop, warp, and layout disruption. Strengths to credit explicitly include the systematic baseline reimplementation under shared noise settings, the CAS residual analysis showing cover-dependent spreading, and the complementary CAS/IAL ablation.","major_comments":[{"comment":"Experimental Settings / Table 2: The headline claim of substantially stronger geometric robustness than eleven baselines is load-bearing on “all models are trained under the same noise settings.” Several END/invertible baselines were originally designed around JPEG/signal layers; under the shared Kornia geometric menu they fall near chance on C&R, shear, rotate, elastic, and jigsaw while CASIAL stays ≥98.44%. Please document capacity/schedule/hyperparameter parity (epochs, LR, noise sampling probabilities, whether spreading variants received equal geometric exposure) and, ideally, report a stronger geometric-augmented baseline protocol (or architecture-appropriate sync auxiliaries). Without this, the margin may partly reflect unequal fitness to this exact attack family rather than a general geometric solution.","section":"Experimental Settings; Table 2"},{"comment":"Motivation & Methods / Abstract: IAL is described as aligning features into a “shared geometry-invariant representation space,” but the evidence is recovery under the same white-box geometric family used in training (crop-resize, erase, shear, rotate, elastic, jigsaw) plus attention realignment on jigsaw (Fig. 5). There is no explicit geometric normalization, group-equivariant design, or test on held-out geometric operators (e.g., perspective/homography, similarity with continuous scale outside the trained range, free-form mesh warps, or print-cam geometry). Please either (i) add such out-of-menu geometric evaluations, or (ii) substantially temper “geometry-invariant” language to “robust under the trained geometric attack family,” and discuss limits of attention-based realignment.","section":"Abstract; Motivation and Methods; Noise layer"},{"comment":"Table 3 / black-box claim: Transfer is demonstrated only on appearance filters (Crayon, Film, Heavy, Layering, Sketch), not on unseen geometric transforms or adaptive/steganalytic attackers. The abstract’s “strong transfer robustness to unseen black-box distortions” is easy to over-read as geometric generalization. Restrict or split the claim (appearance transfer vs. geometric transfer), and add at least one geometric black-box or composition setting (e.g., random perspective + JPEG, or screen-cam-style geometry) so the transfer statement matches the paper’s geometric thesis.","section":"Table 3; Abstract"},{"comment":"CAS formulation (Algorithm S1 / Fig. 2d): Each bit selects a full-resolution candidate F^(0)/F^(1), then L candidates are channel-concatenated to shape (L·C)×H/4×W/4 before IAL fusion. For L=64 this is a large intermediate tensor and a strong inductive bias toward global bit evidence, but the paper does not analyze bit interference, payload scaling, or complexity versus MPB baselines. Please report accuracy vs. L (and wall-clock/memory), and clarify whether performance at L=64 depends on this expensive concatenation rather than on cover-aware selection alone—an ablation with shared low-rank bit embeddings or per-bit lightweight modulators would isolate the claimed mechanism.","section":"CAS Strategy; Algorithm S1"}],"minor_comments":[{"comment":"Fig. 1 is referenced for robustness comparison but the main text does not fully specify what each bar aggregates; align caption with Table 2 protocol (exact strengths: erase 80%, C&R 20%, etc.).","section":"Figure 1"},{"comment":"JND module is important for Table 1 quality yet only lightly specified (“perceptual visibility mask”); cite the exact JND formulation used and whether it is frozen or learned.","section":"Encoder; Supplementary Table S2"},{"comment":"Loss uses MSE on bits (Eq. 2) rather than BCE; a one-sentence justification would help, given binary messages.","section":"Loss Function"},{"comment":"Typos/spacing artifacts appear throughout (e.g., “asasscalingorrotation”, “re-gion removal”, “noiselayersalone”); a careful copy-edit pass is needed.","section":"Introduction; Motivation"},{"comment":"Related Work on transformers is fine but could briefly position against classical sync-template / RST-invariant watermarking to clarify what is new versus classical desynchronization defenses.","section":"Related Work"}],"recommendation":"major_revision","confidential_remarks":"Solid empirical watermarking paper with a real problem focus; not incremental fluff. The main risk for a top venue is over-claiming “geometry-invariant” representations and an apples-to-apples baseline story that may inflate the geometric gap. If the authors add out-of-menu geometric tests and tighten claims, this is likely above threshold; if they only edit prose, I would remain skeptical of the invariance narrative. Scope fit for cs.CV/watermarking venues is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: they name two geometric failure modes (region removal vs desynchronization) and answer them with a concrete message block, CAS, plus spatial attention for aggregation. CAS is the real design move—bits select among cover-derived candidate features rather than upsampling or broadcasting a cover-independent message map. That is a clean alternative to the three MPB patterns they diagram, and Fig. 4 makes the content coupling visible.\n\nEmpirically this is above the usual CV watermarking bar. Eleven baselines retrained under matched resolution, message length, and noise menu; ablations that actually break geometry when CAS or IAL is removed; decoder attention under jigsaw; multi-shape tests; strength curves. Fidelity is not sacrificed for the robustness claim—best PSNR/SSIM/LPIPS in their table while staying ≥98% on heavy crop, erase, shear, rotate, elastic, and jigsaw. Black-box transfer to appearance filters is a nice extra, not the load-bearing result.\n\nThe soft spot is real but bounded. The stress-test note is right that “same noise settings” does not guarantee equal capacity or schedule fairness for older END/invertible pipelines that were not built around geometric layers. Some baselines collapsing near chance may overstate the architectural gap. IAL is standard attention sold as “geometry-invariant representation,” which is stronger language than the evidence: they fit a white-box Kornia menu well and transfer to style filters, not to unseen geometric operators or adaptive/steganalysis attackers. No code or variance reporting yet. None of that falsifies the main empirical claim inside the protocol they actually ran.\n\nWho it is for: people building practical image watermarks for ownership/provenance who still lose under crop and warp. Worth a reading-group slot if that is your lane. I would send it to referees; ask for code, variance, and a fairness check where strong baselines get architecture-appropriate geometric augmentation. Engage with it as a strong systems baseline, not as a theory of invariance.","headline":"Solid systems paper: cover-gated spreading plus attention really does move the geometric numbers, but the headline margin is still partly a matched-noise story.","tokens_in":17733,"tokens_out":513,"would_cite":true,"duration_ms":18717,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Geometric watermark robustness needs global cover-aware message spread plus attention that realigns desynchronized features.","keywords":["image watermarking","geometric robustness","cover-aware message spreading","invariance alignment","spatial attention","desynchronization","region removal","deep learning watermarking"],"falsifier":"Retrain or re-evaluate under stronger or held-out geometric attacks (for example heavier crop ratios, larger rotations, or adaptive desynchronization) and check whether bit accuracy collapses toward chance while baselines do not, or whether a steganalytic/adaptive attacker can strip the CAS residual without destroying the image.","tokens_in":17605,"feed_emoji":"🖼️","tokens_out":885,"duration_ms":16379,"temperature":0.7,"pith_summary":"Deep image watermarks already survive compression and noise, but geometric attacks still break them in two distinct ways: cropping or masking erases the pixels that carried the bits, and rotation, scaling, shear, or scrambling misaligns the remaining evidence so a fixed decoder cannot read it. This paper argues that fixing those failures requires two properties at once—message evidence must be spread across the whole image in a way that depends on the cover content, and the decoder must form geometry-invariant features that can re-aggregate displaced evidence. CASIAL implements the first property with cover-aware message spreading (CAS), which turns each bit into a selector over candidate features derived from the cover, and the second with invariance alignment learning (IAL), spatial attention that reweights reliable regions after distortion. Trained only on differentiable white-box distortions, the system reports near-perfect bit accuracy on six geometric attacks while keeping higher visual quality than eleven prior methods, and it transfers to unseen black-box appearance changes and non-square image shapes.","feed_headline":"Watermarks that survive crop, rotate, and scramble","feed_subtitle":"Cover-aware bit spreading plus attention realignment beat eleven baselines on geometric attacks","key_machinery":"CASIAL’s paired mechanisms: cover image-aware message spreading (CAS), which generates bit-0/bit-1 candidate features from the cover latent and selects them by the message bits before attention fusion, and invariance alignment learning (IAL), spatial attention that maps distorted features into a shared geometry-invariant space.","core_discovery":"Geometric robustness in deep watermarking is not achieved by noise layers alone; it requires explicit global, cover-conditioned message spreading to survive region removal and spatial-attention-based alignment to recover from desynchronization. CASIAL realizes both and, on the authors’ benchmarks, substantially outperforms eleven baselines on six geometric transformations while preserving the best reported visual quality and retaining strong signal, photometric, and black-box transfer performance.","pith_inferences":["The same cover-conditioned spreading idea may transfer to video or audio watermarking where temporal or spectral ‘region removal’ is common.","If IAL’s attention maps are stable under permutation, they could double as a lightweight desynchronization diagnostic or soft sync signal.","Capacity–robustness curves under extreme crop ratios would test whether per-pixel information density from CAS scales beyond the 64-bit setting used here."],"forward_implications":["Message-processing blocks that ignore cover content will remain brittle under heavy crop or erase even if they upsample or broadcast bits.","Spatial attention (or equivalent long-range aggregation) becomes a practical requirement for decoders that must survive rotation, shear, elastic warp, and jigsaw permutation.","JND residual attenuation can raise perceptual quality at matched PSNR by moving energy out of smooth regions, complementary to the robustness gains from CAS and IAL.","A single architecture conditioned on cover latents can serve variable resolutions and aspect ratios without a separate shape-conversion stage."],"fun_headline_variants":["CASIAL spreads bits globally to survive crops and rotations","Cover-aware spreading plus attention beats geometric attacks","Global message spread and invariant alignment toughen watermarks","Watermarks resist region removal via adaptive cover-aware bits","Spatial attention realigns features after scale rotate or scramble"],"cache_read_input_tokens":128,"weakest_assumption_plain":"Training with a fixed menu of differentiable geometric noise layers plus latent attention is assumed sufficient to produce truly geometry-invariant, transferable watermarks without explicit sync marks or adaptive attackers.","fun_headline_variants_meta":{"raw":{"variants":["CASIAL spreads bits globally to survive crops and rotations","Cover-aware spreading plus attention beats geometric attacks","Global message spread and invariant alignment toughen watermarks","Watermarks resist region removal via adaptive cover-aware bits","Spatial attention realigns features after scale rotate or scramble"]},"model":"grok-4.5","effort":"low","cost_usd":0.004381,"raw_usage":{"total_tokens":1350,"prompt_tokens":821,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":43808000,"prompt_tokens_details":{"text_tokens":821,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":463,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":821,"tokens_out":66,"duration_ms":8768,"temperature":1.0,"reasoning_tokens":463,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T22:45:25.870373+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Retrain or re-evaluate under stronger or held-out geometric attacks (for example heavier crop ratios, larger rotations, or adaptive desynchronization) and check whether bit accuracy collapses toward chance while baselines do not, or whether a steganalytic/adaptive attacker can strip the CAS residual without destroying the image.","supporting_citations":[],"review_version":1}