{"id":"59fdde20-62bd-478e-b5c2-7a40535ce6ba","arxiv_id":"2603.14412","paper_version":2,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A multi-scale semi-supervised INR fusion model trained only on a test pair supports arbitrary-scale pansharpening with weight reuse across real remote-sensing scenes.","lead":"G-ZAP fuses a high-resolution panchromatic image with a low-resolution multispectral image using a zero-shot, feature-based implicit neural network that can output any spatial scale. It aims to let one trained model transfer across scenes and sensors without large pretraining or per-image retraining.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Incomplete manuscript blocks verification of the central SOTA and weight-reuse claims that the three-level scheme is said to deliver.","rationale":"The Reader correctly isolates the weakest assumption (sufficiency of multi-scale self-degradation + constructed mid-level supervision for transferable fusion without external data or spatial-fidelity terms) and correctly flags the truncated manuscript as the reason the claim cannot be verified. My stress-test confirms that this is the single most load-bearing concern: every other element of the paper (INR backbone, arbitrary-scale capability, zero-shot framing) is secondary to whether the three-level scheme actually produces the reported SOTA numbers and reusable weights. Because the core equations, architecture, and tables are absent from the provided source, no stronger technical objection (e.g., an algebraic flaw in a loss or an invalid degradation model) can be raised; the evidentiary gap itself is the concern. The recommended verdict therefore remains CONDITIONAL, with no change in direction, pending release of the missing methods/results and preferably public code. The concrete test above is the minimal check that would settle whether the assumption holds.","tokens_in":7624,"tokens_out":663,"duration_ms":6938,"concrete_test":"Obtain the missing method section, architecture diagram, loss equations, and full-resolution quantitative tables (PSNR/SSIM/SAM/ERGAS/QNR or equivalent) for both per-pair and weight-reuse settings on the claimed real-world datasets (WV2/WV3 etc.). Recompute or re-run the weight-reuse column: if the reused-weight metrics fall more than ~5–10 % relative to per-pair retraining or fail to exceed the strongest zero-shot baselines (ZS-Pan, PsDip), the central generalization claim is unsupported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim (Abstract; Contributions 1–3) is that a feature-based INR fusion network trained zero-shot via the three-level multi-scale semi-supervised scheme (Level-0 full-resolution feature preservation, Level-1 4\times with constructed GT, Level-2 multi-scale degradation enabling 16\times relative to Level-0) yields SOTA PAN-scale pansharpening on real data while supporting arbitrary scales and weight reuse competitive with per-pair retraining, without spatial-fidelity constraints. The supplied text only sketches this hierarchy in the Introduction and Related Work; it never supplies the fusion architecture, the precise construction of the mid-level supervised GT, the unsupervised loss terms, the coordinate/feature injection used by the INR, or any quantitative tables/ablations. Consequently the load-bearing premise—that internal multi-scale self-degradation of a single pair is sufficient for transferable spatial–spectral rules under true cross-sensor and full-resolution shift—cannot be audited. The only concrete artifacts present (loss curve in Fig. 8 and qualitative scale-factor visuals in Figs. 9–10) do not address PAN-scale metrics, cross-pair transfer, or the claimed superiority without spatial fidelity. The argument therefore rests on uninspectable experimental support rather than on an internal inconsistency of the described scheme.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes G-ZAP, a zero-shot pansharpening framework that uses a feature-based implicit neural representation (INR) fusion backbone together with a three-level multi-scale semi-supervised training scheme (Level-0 full-resolution feature preservation, Level-1 4\times with constructed mid-level supervision, Level-2 multi-scale degradation enabling relative 16\times learning). The goal is arbitrary-scale HRMS reconstruction from a single PAN–LRMS pair without external pretraining, while also supporting weight reuse across pairs and cross-scene/cross-sensor generalization. The abstract and contribution list claim SOTA PAN-scale results on real-world data, competitiveness of reused weights with per-pair retraining, and that strong performance is obtained without explicit spatial-fidelity constraints.","tokens_in":7969,"tokens_out":989,"duration_ms":10978,"significance":"If the claims hold, the work would be a useful step for practical remote-sensing pansharpening: zero-shot internal learning that still reuses weights, plus continuous-scale output via INR, addresses two real deployment frictions (train–test mismatch and fixed-scale pipelines). The multi-level self-degradation idea is a coherent extension of internal-learning pansharpening (ZS-Pan, PsDip) and feature-based INR fusion (LIIF-style, INF3/FeINFN). The manuscript as supplied, however, does not contain the architecture, losses, GT construction, or quantitative tables needed to verify those claims, so significance remains conditional on a complete experimental package.","major_comments":[{"comment":"The supplied manuscript is incomplete for peer review of the central claims. After Related Work the text jumps to references and supplementary figures (loss curve Fig. 8; qualitative arbitrary-scale Figs. 9–10). There is no method section with the INR fusion architecture, coordinate/feature injection, Level-0/1/2 loss definitions, construction of the mid-level supervised GT, degradation operators/scales, or training schedule. Without these, Contributions 1–3 and the Abstract SOTA / weight-reuse / “no spatial fidelity” claims cannot be audited.","section":null},{"comment":"No quantitative experimental tables or ablations appear in the provided text. Claims of SOTA under PAN-scale fusion, competitiveness of weight reuse vs per-pair retraining, and cross-sensor generalization are therefore unsupported in the document under review. The only concrete artifacts (Fig. 8 convergence; Figs. 9–10 visual comparisons to nearest/bicubic at non-integer scales) do not report standard full-resolution metrics (e.g., QNR, D_λ, D_s) or transfer experiments.","section":null},{"comment":"Introduction framing of Level-1/2 relies on “constructed ground truth” from self-degradation of the test pair. That is standard for internal learning but is load-bearing for the transferability claim. The paper must specify the exact degradation model and show that the resulting supervised terms do not simply encourage identity/upsampling behavior that fails under true cross-sensor or full-resolution shift; this specification is missing from the supplied text.","section":null}],"minor_comments":[{"comment":"Duplicate bibliography entries: Vivone et al. full-scale regression injection appears twice (refs. 32 and 33); Yang et al. Implicit Transformer Network appears twice (refs. 41 and 42).","section":null},{"comment":"Fig. 1 is described in the introduction but not present in the supplied text; ensure the figure and its efficiency/scale axes are included and consistent with the later claims.","section":null},{"comment":"Notation for the three levels (L0/L1/L2) appears only in the supplementary loss curve caption; define symbols consistently when the method section is restored.","section":null},{"comment":"Abstract and contributions assert “first” integration of feature-enhanced INR into zero-shot pansharpening; related work already cites INF3/FeINFN and zero-shot pansharpening—clarify the precise novelty boundary once the method is fully written.","section":null}],"recommendation":"major_revision","confidential_remarks":"The arXiv extract appears truncated (method, experiments, and main figures missing). I am treating this as a major-revision case rather than reject on the assumption that a complete version exists or can be supplied. If the journal receives only this incomplete text, the paper is not reviewable and should be returned. Scope is appropriate for a remote-sensing / CV fusion venue once the experimental package is complete."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The practical pitch is clear: take a feature-based INR fusion backbone, train it zero-shot with a three-level multi-scale semi-supervised scheme (full-res feature preservation, constructed 4× supervision, then further degradations for larger relative scales), and get arbitrary-scale pansharpening plus weight reuse across pairs without external pretraining or explicit spatial-fidelity terms.\n\nWhat is actually new is the packaging, not the ingredients. LIIF-style continuous representation, zero-shot/internal pansharpening (ZS-Pan, PsDip), and multi-scale self-supervision already exist in the cited literature. The contribution is the specific Level-0/1/2 hierarchy aimed at simultaneous arbitrary scale and cross-pair transfer. The problem framing is honest about the real pain points (train-test mismatch, per-image optimization, fixed scale). The qualitative scale-factor figures and the loss curve are consistent with a method that can produce sharper detail than plain upsampling and that converges.\n\nThe soft spot is evidentiary, not conceptual. The supplied text jumps from related work into references and a couple of supplementary figures; the fusion architecture, precise construction of the mid-level “ground truth,” unsupervised loss terms, coordinate/feature injection, quantitative tables, ablations, and cross-sensor transfer numbers are missing. So the load-bearing claim—that internal multi-scale self-degradation of a single pair yields transferable spatial–spectral rules competitive with per-pair retraining and SOTA on real PAN-scale data—cannot be audited from what we have. That is a manuscript-completeness problem, not an internal contradiction in the scheme. Circularity is the usual internal-learning kind (self-generated mid-level targets), not pure tautology.\n\nThis is for people who already work on pansharpening or remote-sensing fusion and care about operational cost (arbitrary scale + weight reuse). It does not reorganize theory. A serious editor should still send a complete version to referees; the idea is concrete and the practical payoff is real if the missing numbers hold. I would not cite it yet, but I would read the full paper and the code if they appear. Engage if you need arbitrary-scale zero-shot fusion; otherwise wait for the complete results.","headline":"Useful packaging of INR + multi-scale internal learning for arbitrary-scale zero-shot pansharpening, but the supplied manuscript is too incomplete to verify the SOTA and weight-reuse claims.","tokens_in":8604,"tokens_out":555,"would_cite":false,"duration_ms":5190,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"G-ZAP learns arbitrary-scale pansharpening from a single test pair and reuses those weights on new scenes without large pretraining or per-image retraining.","keywords":["pansharpening","zero-shot learning","arbitrary scale","implicit neural representation","generalization","image fusion","remote sensing"],"falsifier":"Train once on a single real pair (e.g., WorldView-2), freeze the weights, apply them without fine-tuning to a held-out pair from a different sensor (e.g., WorldView-3 or GaoFen) at full resolution, and compare QNR / spectral–spatial distortion scores against the same model retrained from scratch on that new pair; a large gap would falsify the claimed cross-sensor weight-reuse competitiveness.","tokens_in":8487,"feed_emoji":"🛰️","tokens_out":990,"duration_ms":16014,"temperature":0.7,"pith_summary":"Pansharpening fuses a sharp panchromatic image with a coarser multispectral image to make a sharp multispectral product. Deep models usually need big external training sets and then falter on real full-resolution pairs; prior zero-shot methods fix the distribution gap but must retrain for every pair and stay locked to one fixed scale. This paper claims that a feature-based continuous (INR) fusion network, trained only on the test pair with a three-level multi-scale semi-supervised scheme, can produce high-quality results at any requested scale, transfer its weights across scenes and sensors with little loss, and still match or beat methods that retrain per pair. The practical payoff is efficient real-world use: operators need only the images they already have, can request custom resolutions, and need not re-optimize for every new acquisition. Experiments on multiple real satellite datasets are reported to support state-of-the-art PAN-scale quality and competitive weight reuse.","feed_headline":"Zero-shot pansharpening that reuses weights at any scale","feed_subtitle":"A three-level INR scheme fuses satellite pairs without pretraining or per-image retraining.","key_machinery":"The three-level G-ZAP training scheme on a feature-based INR fusion backbone: Level-0 preserves full-resolution features, Level-1 learns standard 4× upsampling under constructed ground-truth supervision, and Level-2 adds variable-scale degradations so the same continuous decoder also learns a 16× relation back to Level-0, combining unsupervised and constructed supervised losses.","core_discovery":"G-ZAP demonstrates that a feature-enhanced implicit neural representation fusion network, trained zero-shot via a three-level multi-scale semi-supervised loop on a single image pair, achieves state-of-the-art pansharpening at the native PAN scale on real-world data while supporting continuous arbitrary output scales and weight reuse across pairs that remains competitive with full per-pair retraining, without external pretraining or explicit spatial-fidelity constraints.","pith_inferences":["The same multi-scale self-supervision pattern could be tried on related remote-sensing fusion problems (e.g., hyperspectral–multispectral) that still rely on fixed-scale pretrained networks.","If weight reuse holds across sensors, a lightweight edge or on-board model could be updated from one pair and then applied to a stream of later acquisitions.","Continuous INR querying would let analysts request intermediate resolutions for inspection without re-running a full discrete pipeline.","Measuring failure under extreme spectral mismatch or very large scale gaps is a natural next test the paper leaves open."],"forward_implications":["A single trained weight set can be reused on new image pairs instead of retraining for every acquisition.","Output resolution can be chosen continuously, not only at the native PAN size.","Cross-scene and cross-sensor generalization becomes possible from test data alone, without large external corpora.","Prior zero-shot reliance on explicit spatial-fidelity estimation can be dropped under this training scheme.","Deployment cost falls because only the test pair is required and the reported training stabilizes within a few hundred epochs."],"fun_headline_variants":["G-ZAP reuses zero-shot weights for pansharpening at any scale","Feature INR enables arbitrary-scale zero-shot pansharpening","Zero-shot multi-scale training fuses real PAN-MS pairs once","Reusable zero-shot pansharpening matches per-pair quality","Single-pair G-ZAP delivers PAN-scale fusion across sensors"],"cache_read_input_tokens":128,"weakest_assumption_plain":"Multi-scale self-degradation of one test pair plus constructed mid-level labels is assumed enough for the network to learn fusion rules that stay accurate under true cross-sensor and full-resolution shift without external data or spatial-fidelity terms.","fun_headline_variants_meta":{"raw":{"variants":["G-ZAP reuses zero-shot weights for pansharpening at any scale","Feature INR enables arbitrary-scale zero-shot pansharpening","Zero-shot multi-scale training fuses real PAN-MS pairs once","Reusable zero-shot pansharpening matches per-pair quality","Single-pair G-ZAP delivers PAN-scale fusion across sensors"]},"model":"grok-4.5","effort":"low","cost_usd":0.005514,"raw_usage":{"total_tokens":1491,"prompt_tokens":810,"num_sources_used":0,"completion_tokens":101,"cost_in_usd_ticks":55140000,"prompt_tokens_details":{"text_tokens":810,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":580,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":810,"tokens_out":101,"duration_ms":4856,"temperature":1.0,"reasoning_tokens":580,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T21:16:13.800551+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train once on a single real pair (e.g., WorldView-2), freeze the weights, apply them without fine-tuning to a held-out pair from a different sensor (e.g., WorldView-3 or GaoFen) at full resolution, and compare QNR / spectral–spatial distortion scores against the same model retrained from scratch on that new pair; a large gap would falsify the claimed cross-sensor weight-reuse competitiveness.","supporting_citations":[],"review_version":1}