{"id":"138988fb-10dd-49cb-8826-5e0c48111227","arxiv_id":"2508.16284","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"A hybrid CNN-transformer with auxiliary noiseprint features achieves competitive forgery detection and localization on ID documents.","lead":"EdgeDoc is a hybrid CNN-transformer model that detects and localizes forged regions in ID documents by combining image features with noiseprint traces. It placed third in the ICCV 2025 DeepID Challenge and reportedly beats baseline approaches on the FantasyID dataset.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Real-world effectiveness claim hinges on FantasyID representativeness; unvalidated dataset-to-deployment generalization is the weakest link.","rationale":"The reader correctly identified the dataset representativeness assumption as the weakest link. My stress-test agrees: the central claim's real-world validity depends entirely on FantasyID being a faithful proxy for actual forgeries, and the noiseprint component creates a specific risk that the benchmark is skewed toward manipulations that produce detectable fingerprint anomalies. The corrupted full text prevents verifying the experimental setup, so no additional evidence can be weighed; this supports the reader's UNVERDICTED verdict. However, the concern is not a fatal flaw—it is a call for external validation. I therefore recommend no change to the verdict: the paper remains unverdictable based on the available abstract, and the specific concern I raise would require access to the dataset or additional experiments. The agreement is 'agree' because the reader's weakest_assumption matches my own, though I add the noiseprint-based channel as a concrete mechanism for why the benchmark could be biased.","tokens_in":15046,"tokens_out":2943,"duration_ms":34909,"concrete_test":"Obtain the FantasyID dataset and the DeepID Challenge evaluation protocol, then run a cross-dataset generalization test: train/validate EdgeDoc and the same baselines on FantasyID, and test on a real-world ID forgery benchmark such as MIDV-2020 or a newly collected set of tampered ID documents containing both classical and AI-based manipulations. If EdgeDoc's relative ranking reverses or its performance drops disproportionately (e.g., >10% mIoU/AP) compared to the FantasyID results, the 'realworld effectiveness' claim is invalid. Additionally, inspect FantasyID's manipulation generation pipeline to determine whether forgeries include diffusion-based edits; if not, the dataset lacks coverage of a major contemporary threat.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that EdgeDoc is competitive and effective—rests on a single evaluation: outperforming baselines on the FantasyID dataset and placing third in the ICCV 2025 DeepID Challenge. The abstract then generalizes to 'realworld scenarios.' For this empirical claim to hold, FantasyID must be representative of actual KYC document forgeries. This is the weakest assumption. EdgeDoc explicitly uses noiseprint features, which encode camera-sensor and editing fingerprints. If FantasyID forgeries are predominantly classical manipulations (splicing, copy-move, inpainting) that leave strong noiseprint traces, the method's superiority may be an artifact of cue leakage rather than a general detection capability. Real-world forgeries increasingly involve generative AI (GANs, diffusion) that do not produce traditional noiseprint anomalies. The full text is corrupted in this submission, so FantasyID's construction (synthetic vs. real, manipulation types, document variety, baseline choices, and metric definitions) cannot be independently verified. The claim that results 'highlight effectiveness in realworld scenarios' is therefore not supported by the evidence visible in the abstract alone.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes EdgeDoc, a hybrid CNN-transformer model augmented with auxiliary noiseprint features for detecting and localizing forgeries in ID documents. The abstract reports that EdgeDoc placed third in the ICCV 2025 DeepID Challenge and outperforms baseline methods on the FantasyID dataset, interpreting this as evidence of effectiveness in real-world KYC scenarios. The provided manuscript body is unreadable due to widespread encoding corruption; only the abstract is intelligible. Consequently, the architecture, the noiseprint feature formulation, the training and evaluation protocols, the dataset description, and the quantitative results are not available for verification.","tokens_in":15307,"tokens_out":3563,"duration_ms":39456,"significance":"If substantiated, the paper would offer a useful, competitive approach for document forgery detection: the combination of a lightweight convolutional transformer with noiseprint features is a sensible direction, and an external challenge placement provides some signal of practical competitiveness. However, the significance depends entirely on the evaluation's validity and generalizability. Because the body text is corrupted and the abstract contains no numbers, no quantitative claim can currently be checked. The paper's central claim is plausible but unverified at this stage.","major_comments":[{"comment":"The manuscript body is unreadable: it consists of mojibake with no parseable sentences, equations, tables, or figure captions. As a result, the architecture details, the noiseprint feature computation, the training procedure, dataset statistics, baseline definitions, and metric definitions are all missing. The central claim of outperforming baselines on FantasyID cannot be verified from the supplied file. This is the load-bearing evidence for the paper and must be addressed by providing a readable manuscript.","section":"Full text (all body sections)"},{"comment":"The abstract reports 'third place' and 'outperforms baseline approaches' but gives no numerical values, no metrics (e.g., F1, IoU, AUC), no error bars, and no comparison protocol. Even if the body were readable, the abstract alone does not support the effectiveness claim. The authors should state the quantitative results and evaluation metrics in the abstract or, at minimum, clearly in the results section.","section":"Abstract, paragraph 3"},{"comment":"The claim that results 'highlight its effectiveness in realworld scenarios' generalizes from a single dataset (FantasyID). There is no visible evidence that FantasyID is representative of real-world KYC forgeries or of the manipulation types encountered in practice. The manuscript should discuss the dataset's construction (synthetic vs. real, manipulation types, document types) and should address whether noiseprint-based features are sufficient for modern generative manipulations. Without this, the real-world generalization claim is unsupported.","section":"Abstract, paragraph 3 / Results (missing)"}],"minor_comments":[{"comment":"The project page URL contains a space: 'https://www.idiap. ch/paper/edgedoc/'. Also, 'realworld' should be 'real-world'.","section":"Abstract, last line"},{"comment":"The FantasyID dataset and the ICCV 2025 DeepID Challenge are not cited with full references in the visible text. Provide citations and, if available, dataset documentation.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The supplied PDF appears to have a serious text-extraction corruption issue. Before substantive review can proceed, the authors should be asked to submit a readable version. The topic is within the journal's scope, but the current file is not reviewable in its present form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honest take: this is a standard, workmanlike application paper with a real external validation signal. The new bit is the specific combination of a lightweight CNN-transformer with noiseprint features applied to ID-document forgery detection and localization. Both ingredients are known in forensics, so the contribution is bounded, but a new combination targeting a real KYC problem can still be a legitimate contribution. Third place in the ICCV 2025 DeepID Challenge is a concrete datum, and the FantasyID results, if they match the abstract, put the method above its baselines.\n\nI could not inspect the body—the full text in this submission is corrupted (all replacement characters), so my read is restricted to the abstract. That means I cannot verify the architecture details, metric definitions, baseline choices, or whether the numbers come with error bars or ablations. The central claim as stated is modest: competitive rather than SOTA. So I would not call the paper unsound from what is visible.\n\nThe soft spots, in proportion: the abstract's last sentence—\"highlighting its effectiveness in realworld scenarios\"—overreaches relative to the evidence shown. One challenge and one dataset are not real-world deployment. If FantasyID is synthetic or dominated by classic copy-move/splicing manipulations, that is exactly where noiseprint features are strongest, and the method may not generalize to GAN- or diffusion-based forgeries, which are the emerging threat in KYC fraud. That concern is real but it is about the evidence, not necessarily the method. A good reviewer should ask the authors to report per-manipulation-type results and, ideally, evaluate on a second dataset. The \"novel approach\" phrasing in the abstract also oversells what is an incremental combination, but that is a minor framing issue.\n\nWho is this for? People building document-forensics pipelines or working on the DeepID challenge line. They will get a useful baseline architecture recipe and a competitive point of reference. It does not deserve a desk reject; it deserves a normal peer review with attention to dataset representativeness and the real-world generalization language. I would not cite it in my own work before seeing the body, but I would not be surprised if it holds up.","headline":"A modest, credible application paper with a concrete challenge result; the real uncertainty is whether the FantasyID evaluation supports the real-world claim, and the corrupted full text makes that unanswerable from this submission.","tokens_in":15728,"tokens_out":2279,"would_cite":false,"duration_ms":26031,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"EdgeDoc claims a hybrid CNN-transformer with noiseprint features detects and localizes forged ID documents better than baselines on the FantasyID benchmark.","keywords":["document forgery detection","forgery localization","ID documents","noiseprint","CNN-transformer hybrid","KYC security","image forensics"],"falsifier":"Construct a test set of forged ID documents that have been edited and then printed, re-scanned, or saved with strong JPEG/WebP recompression—manipulations common in real fraud attempts. If EdgeDoc's detection accuracy on this set falls to the level of a plain CNN baseline, the central claim that noiseprint-augmented hybrid features are decisive would be falsified.","tokens_in":14971,"feed_emoji":"🪪","tokens_out":5075,"duration_ms":51291,"temperature":0.7,"pith_summary":"The paper aims to show that a hybrid architecture—a lightweight convolutional transformer combined with auxiliary noiseprint features—can both flag a forged identity document and show where it was altered. This matters because KYC and remote-onboarding checks need to catch subtle edits that look natural to human reviewers, and localization gives investigators a concrete region to inspect. The evidence is empirical: EdgeDoc placed third in the ICCV 2025 DeepID Challenge and outperformed baseline approaches on the FantasyID dataset. If correct, the result points to a practical recipe for document forgery detection that fuses low-level sensor-noise cues with context-aware transformer processing.","feed_headline":"Hybrid model spots forged IDs and localizes tampering","feed_subtitle":"A CNN-transformer with noiseprint cues outperforms baselines on FantasyID, strengthening KYC checks.","key_machinery":"The load-bearing component is the EdgeDoc architecture, which fuses two input streams: the document image processed by a lightweight convolutional transformer, and auxiliary noiseprint features. Noiseprint is a per-image fingerprint of the sensor and processing pipeline that produced the image; any inserted or reworked region has an inconsistent fingerprint, so the noiseprint stream gives the transformer a localization signal aligned with tampered pixels. The design lets the transformer's self-attention combine layout and text semantics with these low-level forensics cues in one forward pass.","core_discovery":"On its own terms, this paper's claim is that fusing auxiliary noiseprint features into a lightweight convolutional transformer yields a single model that both detects whether an ID document has been forged and localizes the altered region. The paper reports that the resulting system, EdgeDoc, placed third in the ICCV 2025 DeepID Challenge and outperformed the baseline approaches it was compared with on the FantasyID dataset. The stated reason the combination works is that the transformer captures document structure and context while the noiseprint stream exposes the traces left by splicing or re-editing, traces that are hard to see in pixel space alone.","pith_inferences":["The gain from noiseprint is likely to shrink when forgeries are printed, re-scanned, or aggressively recompressed, because those operations destroy the sensor-noise footprint; the paper does not test this regime.","Because ID documents are highly standardized layouts, the same twin-stream architecture should transfer to passports, driver's licenses, visas, and diplomas with little change.","An ablation that removes the noiseprint branch on FantasyID would isolate how much of EdgeDoc's edge comes from the auxiliary features versus the transformer backbone; the paper does not report this split explicitly."],"forward_implications":["KYC and remote-onboarding providers could use EdgeDoc to flag suspect documents and automatically highlight the specific forged region for manual review.","Adding a noiseprint stream is a general recipe that strengthens a lightweight transformer for forgery detection without a separate forensics classifier.","Because detection and localization happen in one model, deployment cost stays close to a single backbone instead of an ensemble of detectors and segmenters.","The DeepID Challenge ranking is evidence that the method generalizes beyond its training distribution at least as well as the specific baselines it was compared with."],"supporting_citations":[],"fun_headline_variants":["Hybrid model with noiseprint detects and localizes forged IDs","EdgeDoc hybrid flags and localizes ID tampering","Noiseprint-aware transformer catches forged ID documents","EdgeDoc fuses noiseprint to pinpoint forged ID regions","Hybrid CNN-transformer uses noiseprint to spot ID fakes"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The paper's claim that the method works in real-world scenarios depends on the FantasyID dataset and the DeepID Challenge evaluation representing the forgeries that actually occur in KYC fraud; if that benchmark is narrow or synthetic, the reported advantage may not carry over.","fun_headline_variants_meta":{"raw":{"variants":["Hybrid model with noiseprint detects and localizes forged IDs","EdgeDoc hybrid flags and localizes ID tampering","Noiseprint-aware transformer catches forged ID documents","EdgeDoc fuses noiseprint to pinpoint forged ID regions","Hybrid CNN-transformer uses noiseprint to spot ID fakes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000473,"raw_usage":{"total_tokens":2141,"prompt_tokens":656,"completion_tokens":1485,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":400,"completion_tokens_details":{"reasoning_tokens":1397}},"tokens_in":400,"tokens_out":1485,"duration_ms":13119,"temperature":1.0,"reasoning_tokens":1397,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:22:36.081071+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a test set of forged ID documents that have been edited and then printed, re-scanned, or saved with strong JPEG/WebP recompression—manipulations common in real fraud attempts. If EdgeDoc's detection accuracy on this set falls to the level of a plain CNN baseline, the central claim that noiseprint-augmented hybrid features are decisive would be falsified.","supporting_citations":[],"review_version":1}