{"id":"dc31a1d0-1fc7-4c99-bdf3-65c0ddc360e6","arxiv_id":"2507.20404","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The second ID-card presentation attack detection competition produced winners with 11.34% and 6.36% equal error rates on the shared-data and open-data tracks, improving on the first edition's 21.87% EER.","lead":"This paper reports the results of the second competition on detecting fake ID cards in photos, with two tracks and 74 submitted models. The best systems reached equal error rates near 11% and 6%, showing progress but persistent difficulty on some countries and attack types.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Track 2 winner Incode's 0.04-point AVRank margin over Baseline-1 is within sampling error; the decisive BPCER100 advantage is 30 errors in 5,000 bona-fide images, so the winner claim is not statistically supported without confidence intervals or weight sensitivity analysis.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the Track 2 winner margin is within noise and the winner could change under different weighting. My analysis decomposes the 0.04-point AVRank difference and confirms that the decisive BPCER100 advantage is roughly 30 errors out of 5,000 bona-fide images, well inside sampling uncertainty, while BPCER20 and EER favor Baseline-1. The paper provides no uncertainty estimates and no sensitivity analysis for Eq. 3, and the reporting procedure (best submission per participant, unchanged 2024 test set) further weakens the precision of the winner claim. Track 1, by contrast, has a larger margin (40.48 vs 43.62) and is not threatened by this concern. The paper is a useful competition report and its tables are transparent, but the central winner claim for Track 2 should be stated conditionally. This does not move the reader's CONDITIONAL verdict, so the recommended adjustment is UNCHANGED.","tokens_in":12467,"tokens_out":6342,"duration_ms":74913,"concrete_test":"Obtain per-image scores from the evaluation server for Incode and Baseline-1 on the 23,851-image test set. Perform a paired bootstrap over images (10,000 resamples, stratified by country and PAIS), computing for each resample BPCER10/20/100, EER, and AVRank under Eq. 3, and report the two-sided 95% confidence interval for the AVRank difference. If the interval includes 0, as expected from the 30/5,000 BPCER100 gap, the 'Incode wins' conclusion is unsupported. As a complementary check, recompute the leaderboard with equal weights (1/3 each) and with alternative weights corresponding to plausible cost ratios; if Baseline-1 wins under any reasonable weighting, the paper should state that the ranking is weight-dependent rather than a unique winner.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that Incode wins Track 2 with AVRank 14.76% (Tables 6/7, Eq. 3). The winning margin over Baseline-1 is 0.04 AVRank points, and it is not driven by any robust component. In weighted terms: Incode gains 0.2*(3.06-2.56)=+0.100 on BPCER10 and 0.5*(23.64-23.04)=+0.300 on BPCER100, but Baseline-1 gains 0.3*(9.08-7.90)=+0.354 on BPCER20. The BPCER100 advantage is 23.04% vs 23.64%, i.e., about 30 errors out of 5,000 bona-fide images, which is within sampling noise; the paper provides no confidence intervals, bootstrap, or other uncertainty estimates. With equal weights on the three BPCER thresholds, Baseline-1 (11.533%) would beat Incode (11.560%), so the ranking can be reversed by a plausible change in the hand-selected weights. Section 3 reports only the best submission per participant, and Section 2.3 reuses the unchanged 2024 sequestered test set, so the 0.04-point margin is also the result of best-of-N selection on a fixed benchmark. The abstract's EER for Dragons (11.44%) disagrees with Table 6 (11.34%), but this is minor; the real weak point is that the Track 2 winner is not distinguishable from a baseline that has a lower EER and better BPCER20.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports the organization and results of the second Presentation Attack Detection competition on ID cards (PAD-ID Card 2025). Track 1 provides a shared synthetic training set of 12,000 images; Track 2 allows participants to use any open or private data. All submissions are evaluated on a sequestered test set of 23,851 images from four countries using ISO/IEC 30107-3 metrics, with the winner determined by AVRank, a weighted combination of BPCER10, BPCER20, and BPCER100. The reported winners are Dragons (Track 1, AVRank 40.48%) and Incode (Track 2, AVRank 14.76%), with Incode improving on the 2024 competition's best AVRank of 74.30%. The paper also describes the dataset generation, baselines, participant methods, and DET curves.","tokens_in":12831,"tokens_out":4372,"duration_ms":47104,"significance":"The main contribution is an independent, standardized benchmark of current ID-card PAD algorithms on a sequestered test set, together with a shared synthetic training dataset and an automated evaluation platform. These assets are valuable to the community, and the reported improvement over the first edition suggests genuine progress in the field. The paper is also honest about limitations, explicitly noting in Section 7 that model efficiency and the cost of inference need consideration in future editions. However, the significance of the headline claims is conditional on the statistical robustness of the winner ranking, which the paper currently does not establish.","major_comments":[{"comment":"The Track 2 winner claim is not statistically supported. Incode's AVRank (14.76%) exceeds Baseline-1 (14.80%) by only 0.04 percentage points, while Baseline-1 has a lower EER (6.07% vs. 6.36%) and a lower BPCER20 (7.90% vs. 9.08%). The decisive BPCER100 advantage (23.04% vs. 23.64%) corresponds to roughly 30 of the 5,000 bona fide test images. Without confidence intervals, bootstrap estimates, or a sensitivity analysis over the AVRank weights in Eq. (3), the reported winner ranking cannot be distinguished from noise; an equal-weight version of Eq. (3) would rank Baseline-1 ahead (11.533% vs. 11.560%). The paper should report uncertainty estimates and show that the winner determination is robust to reasonable changes in the metric weights.","section":"§4, Eq. (3); Table 7"},{"comment":"The protocol's best-of-N selection and the unchanged test set weaken the comparative claims. Section 3 states that 'only the best submission per participant is shown and reported,' and Section 2.3 states that the sequestered test set 'remains unchanged from the previous edition.' Since participants could submit multiple models, the reported margins—particularly the 0.04-point Track 2 margin—reflect selection over submissions as well as algorithm quality. The reuse of the 2024 test set also allows participants to have tuned to the exact evaluation distribution. The paper should report all submissions or at least the number of submissions per team and discuss the effect of this selection on the reported results.","section":"§3; §2.3"}],"minor_comments":[{"comment":"The abstract states an EER of 11.44% for the Track 1 winner, while Table 6 reports 11.34%; please correct the inconsistency and ensure that all reported numbers match between the abstract, tables, and text.","section":"Abstract; Table 6"},{"comment":"In Table 3, the Validation count for Composite is written as '15.900'; this should be '15,900' to match the comma format used in the other entries.","section":"Table 3"},{"comment":"The IDCH description contains a typo: 'boan fide' should be 'bona fide'.","section":"§5.3.5"},{"comment":"The captions for Figures 2 and 3 refer to per-country panels '(e),(f), and (g)', but each figure has eight panels (a)–(h); please update the caption text to reference all per-country panels.","section":"Figure 2; Figure 3"},{"comment":"Section 7 raises model efficiency as a concern for deployment but provides no runtime data; either add timing results or explicitly mark this as an open direction for future work.","section":"§7"}],"recommendation":"major_revision","confidential_remarks":"The paper is a competition report with a clear empirical contribution. The main concern is that the Track 2 winner is separated from the top baseline by a margin that is almost certainly within sampling variability, and the paper provides no uncertainty quantification. The editor may wish to emphasize to the authors that the headline claim of a Track 2 winner should be either supported with confidence intervals and weight sensitivity analysis or toned down to a statement about a nominal ranking. The rest of the manuscript, including the dataset description and the comparison across teams, is useful and appropriate for the venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful second-edition competition report, not a methods paper. The new assets are the shared synthetic training set, the two-track design, and the automatic EvalAI-based evaluation platform. The longitudinal comparison is the main value: Track 2's winner improves on the first edition's best AVRank from 74.30% to 14.76%, and EER from 21.87% to 6.36%, which is a big jump regardless of who gets the trophy.\n\nThe tables and equations line up; I checked the AVRank arithmetic. One minor inconsistency: the abstract says Track 1 winner EER 11.44%, Table 6 says 11.34%, so that should be fixed.\n\nThe soft spot is the Track 2 winner claim. Incode beats Baseline-1 by 0.04 AVRank points (14.76 vs 14.80). The margin comes from better BPCER10 and BPCER100, but worse BPCER20. On the BPCER100 component, the gap is 23.04% vs 23.64% on 5,000 bona fide images, i.e., about 30 images. No confidence intervals are reported, and with equal weights on the three BPCER terms Baseline-1 would edge ahead. So the paper should either report uncertainty (bootstrap, Clopper-Pearson intervals) or soften the winner language for Track 2. The abstract's 'Incode reached the best results' is stronger than the evidence supports.\n\nThe best-of-N reporting (only best submission per team shown) is standard for competitions, but it adds selection bias on top of a fixed test set. Reusing the 2024 test set is a double-edged sword: it enables clean year-over-year comparison, but repeated evaluation against the same images erodes the benchmark's strength. The paper should acknowledge this.\n\nWhere the paper is solid: the dataset description is concrete, the baselines are sensible, the results tables are complete, and the organizers shipped a useful artifact—the synthetic dataset and the evaluation platform. That deserves credit.\n\nBottom line: worth a serious referee. The central longitudinal improvement is robust even if the specific Track 2 winner is not statistically distinguishable from Baseline-1. I'd ask for a sensitivity analysis on the AVRank weights and some error bars before accepting, but this is a conditional accept, not a reject.","headline":"Useful second-edition competition report with real assets (synthetic training set, two-track protocol, longitudinal benchmark), but the Track 2 winner claim rests on a 0.04-point margin that needs uncertainty analysis before it can be taken at face value.","tokens_in":13488,"tokens_out":2364,"would_cite":true,"duration_ms":26126,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports the outcome of the second open competition on ID-card presentation attack detection, claiming that the field improved sharply: the best open-track system reached AVRank 14.76% and EER 6.36%, versus 74.30% and 21.87% in…","keywords":["presentation attack detection","ID card","biometrics","benchmark","AVRank","sequestered test set","synthetic dataset","spoof detection"],"falsifier":"If the organizers make per-image scores public, a bootstrap test on AVRank would settle the headline claim: if the 95% confidence intervals for Incode and Baseline-1 overlap, the Track 2 win is not statistically distinguishable under the chosen metric; an even simpler check is to recompute the leaderboard with equal weights on the three BPCER points and see whether the ranking flips.","tokens_in":12282,"feed_emoji":"🪪","tokens_out":9734,"duration_ms":104110,"temperature":0.7,"pith_summary":"This paper reports the outcome of the second open competition on detecting presentation attacks on ID cards, meaning attempts to pass off printed, displayed, or altered copies of identity documents as genuine. Its central claim is that the field is improving: on a sequestered test set of 23,851 images, the best submitted system in the open-data track reached an average ranking of 14.76% and an equal-error rate of 6.36%, compared with 74.30% and 21.87% for the best system in the first edition. The paper also introduces a shared synthetic training set and an automatic evaluation platform, and documents that the hardest remaining cases are screen attacks in the open track and unseen-country generalization, especially Panama. A sympathetic reader takes away a concrete benchmark number to beat, not a proof that any particular architecture is best.","feed_headline":"ID-card spoof score drops to 14.76% from 74.30% in new test","feed_subtitle":"Track 2 winner also cut equal-error rate to 6.36%, from 21.87% in 2024, on a 23,851-image test set.","key_machinery":"The load-bearing mechanism is the evaluation protocol: a sequestered test set of 23,851 images from Chile, Guatemala, Panama, and Mexico, with bona fide samples plus screen, print, and composite attacks; an automated evaluation platform that scores each submission through a Docker API; and a headline metric AVRank = 0.2·BPCER10 + 0.3·BPCER20 + 0.5·BPCER100, which weights the most security-critical operating points most heavily. The paper also provides baselines, an EfficientNetV2-S trained on the shared synthetic data for Track 1 and three MobileViTv2 baselines trained on private and open data for Track 2, so that submitted models are compared against reproducible reference points.","core_discovery":"The discovery is the measured state of ID-card presentation attack detection in 2025, established by evaluating 74 submitted models on a fixed sequestered test set from four countries. In Track 1, where all teams trained on the shared synthetic dataset, the Dragons team won with AVRank 40.48% and EER 11.34%, using a CLIP visual encoder fine-tuned with LoRA and a YOLOv8-based document crop. In Track 2, where teams could use any data, Incode won with AVRank 14.76% and EER 6.36%, using an ensemble of two CNNs and one transformer trained on roughly 218,000 proprietary images. These numbers improve on the first edition's best AVRank of 74.30% and EER of 21.87%, and the paper interprets the gap as evidence that larger stores of bona fide images, not only attack diversity, drive generalization.","pith_inferences":["If the organizers release per-image scores, the same leaderboard could be recomputed under alternative weighting schemes, turning the competition's single ranking into a sensitivity analysis; under some weightings the Track 2 order may differ.","A direct next experiment is to train the Track 1 shared synthetic dataset with the Track 2 winner's ensemble recipe; the comparison would isolate how much of the two-track gap comes from data scale versus architecture.","The paper's emphasis on bona fide image count suggests a cheap test: add synthetic bona fide variations of real card templates and measure whether AVRank on Panama and screen attacks drops.","The Track 1 result also suggests that contrastive language-image pretraining can serve as a data-efficient starting point for ID-card PAD; replacing CLIP with a CNN or supervised ViT in the same pipeline would test that hypothesis."],"forward_implications":["Future submissions can use AVRank 40.48% and 14.76% as the published state of the art for the two tracks on this test set.","The Track 1 winner's approach, segmentation-based cropping plus a fine-tuned CLIP encoder, shows that pretrained visual-language features transfer to unseen ID-card countries when training data is limited.","The Track 2 result ties performance mainly to training data scale and diversity: the winning ensemble used about 218,000 images and beat all three baselines, while smaller-data submissions remained far behind.","The organizers' call for future editions to track inference time recognizes that several foundation-model submissions take 1–2 minutes per image, too slow for online onboarding."],"supporting_citations":[{"why":"Supplies the sequestered test set and the first-edition results (74.30% AVRank, 21.87% EER) that all 2025 improvements are measured against.","marker":"[14]"},{"why":"Provides the evaluation platform used to run every submission and compute the leaderboard metrics.","marker":"[18]"},{"why":"Supplies the ImageNet pretrained weights that initialize the Track 1 baseline and most submitted models.","marker":"[4]"},{"why":"Defines the EfficientNetV2-S architecture used as the Track 1 baseline.","marker":"[13]"},{"why":"Cited as the MobileViTv2 architecture used for the three Track 2 baselines.","marker":"[6]"},{"why":"Provides the CLIP model that the Track 1 winner fine-tuned with LoRA to reach first place.","marker":"[3]"},{"why":"Describes the low-rank adaptation technique used to fine-tune the winner's CLIP encoder.","marker":"[16]"}],"fun_headline_variants":["ID-card spoof error drops fivefold in new competition","Best ID-card PAD score improves to 14.76% from 74.30%","Second ID-card spoof test: error rate falls to 14.76%","74-model ID-card spoof test: winner hits 14.76%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ranking of winners rests on the competition's hand-chosen weighted average of error rates at three operating points, and a different weighting, or any uncertainty estimate, could reorder the top teams, especially in Track 2 where the winner's margin over a baseline is 0.04 points.","fun_headline_variants_meta":{"raw":{"variants":["ID-card spoof error drops fivefold in new competition","Best ID-card PAD score improves to 14.76% from 74.30%","Second ID-card spoof test: error rate falls to 14.76%","74-model ID-card spoof test: winner hits 14.76%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000607,"raw_usage":{"total_tokens":2855,"prompt_tokens":996,"completion_tokens":1859,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":612,"completion_tokens_details":{"reasoning_tokens":1775}},"tokens_in":612,"tokens_out":1859,"duration_ms":20252,"temperature":1.0,"reasoning_tokens":1775,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T13:34:14.010784+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If the organizers make per-image scores public, a bootstrap test on AVRank would settle the headline claim: if the 95% confidence intervals for Incode and Baseline-1 overlap, the Track 2 win is not statistically distinguishable under the chosen metric; an even simpler check is to recompute the leaderboard with equal weights on the three BPCER points and see whether the ranking flips.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the sequestered test set and the first-edition results (74.30% AVRank, 21.87% EER) that all 2025 improvements are measured against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the ImageNet pretrained weights that initialize the Track 1 baseline and most submitted models."},{"cited_title":"Tan and Q","cited_arxiv_id":null,"evidence_quote":"Defines the EfficientNetV2-S architecture used as the Track 1 baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Cited as the MobileViTv2 architecture used for the three Track 2 baselines."},{"cited_title":"Cherti, R","cited_arxiv_id":null,"evidence_quote":"Provides the CLIP model that the Track 1 winner fine-tuned with LoRA to reach first place."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the low-rank adaptation technique used to fine-tune the winner's CLIP encoder."}],"review_version":1}