{"id":"b476a101-1523-4e70-a147-ea22a6832796","arxiv_id":"2501.19325","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A hybrid deep-learning and genetic-algorithm solver reports state-of-the-art accuracy on large square-tile puzzles, including Portuguese tile panels and highly eroded images.","lead":"An Israeli-Portuguese team combines a deep-learning compatibility score with a genetic algorithm to reassemble square-tile jigsaw puzzles, reporting near-perfect reconstruction of Portuguese azulejo panels and large eroded puzzles. The work is a candidate practical tool for archaeology, art restoration, and document forensics, where manual reassembly is extremely slow.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Portuguese-tile SOTA may rest on unverified train/test overlap: the 208 online training images (some tourist photos) are not checked for duplicates of the 24 MNAz test panels, so reported 95.2%/89.4% could reflect memorized piece pairs.","rationale":"The reader identified cropping misalignment as the weakest assumption. I agree that automated cropping could inflate the Portuguese-tile numbers, but the more fundamental risk is unverified train/test overlap: the training source is Internet images of the same museum collection, and no duplicate check is mentioned. This is not an accusation of misconduct; it is an unguarded assumption that should be tested before the central SOTA claim is accepted. The GA best-of-50 reporting is also a concern, but Table III shows the average accuracy still far exceeds prior methods, so that issue is less decisive. Data leakage, by contrast, would directly collapse the headline result. I therefore retain the reader's CONDITIONAL verdict and propose a concrete deduplication test as the key condition for acceptance.","tokens_in":19430,"tokens_out":8263,"duration_ms":87500,"concrete_test":"Compute robust perceptual hashes (e.g., pHash/dHash) of the 24 MNAz test panels and the 208 online training images, then manually inspect near-duplicates allowing for different crops, scaling, and lighting. Also check the 9 validation images. If any test panel appears in training, retrain the DLCM on the deduplicated set and re-run the Portuguese-tile experiments; if the Top-1 and neighbor-accuracy numbers in Tables I and II drop materially, the SOTA claim is invalidated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section V-A states that the 24 MNAz test panels were excluded from CNN training, but the training set was built from 208 images gathered from the Internet, including images 'taken by casual tourists,' and no deduplication against the test set is described. This matters because the DLCM is trained on concatenated piece pairs from each training image; if a test panel's pieces appear in training, the network can memorize specific pairwise adjacencies rather than learn a generic compatibility measure. That would directly inflate the Top-1 scores in Table I and the neighbor-accuracy results in Table II (95.2% Type-1 and 89.4% Type-2 with known dimensions), which are the paper's headline SOTA evidence. The cropping-alignment limitation is acknowledged and described as rare; the overlap risk is unaddressed and, if real, would invalidate the central claim. This is a load-bearing assumption because the entire Portuguese-tile SOTA claim depends on clean held-out generalization.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-stage framework for square-piece visual reconstruction: a deep CNN-based compatibility measure (DLCM) trained on full piece pairs via binary cross-entropy, followed by a genetic algorithm solver with a hierarchical crossover. The framework is evaluated on Portuguese tile panels (24 MNAz test panels), synthetic JPP benchmarks, eroded-boundary puzzles, and shredded documents. The authors report state-of-the-art results, including 95.2% and 89.4% neighbor accuracy for Type-1 and Type-2 Portuguese tile panels with known dimensions (Table II), and 16.2%/35.1% average improvements over Bridger et al. on eroded puzzles (Table VI). The contribution also includes a new benchmark dataset of Portuguese tiles.","tokens_in":19624,"tokens_out":6310,"duration_ms":58275,"significance":"If the reported numbers are reproducible under clean held-out conditions, the paper demonstrates a substantial empirical advance, particularly for Type-2 puzzles and eroded boundaries, and the ablation study (Table III) is a useful decomposition of the GA phases. The release of the Portuguese tile benchmark is a valuable asset to the community. However, the central SOTA claims are currently supported by comparisons that are vulnerable to data-contamination risk and to best-of-N reporting bias; the paper provides no statistical evidence for the advantage over baselines. These issues must be resolved before the claims can be accepted.","major_comments":[{"comment":"Section V-A states that the 24 MNAz test panels were excluded from CNN training, but the 208 Internet-acquired training images (some 'taken by casual tourists') are not checked for duplication or near-duplication with the test panels. Because the DLCM is trained on concatenated piece pairs, a tourist photo containing any test panel would provide direct supervision for exactly the adjacency pairs used in the Top-1 evaluation of Table I and the neighbor-accuracy evaluation of Table II. The reported 95.2% and 89.4% accuracies therefore rest on an unverified assumption of clean held-out generalization. The authors should verify, by image retrieval or manual inspection, that no test panel appears in the training images, or retrain on a deduplicated training set and report the resulting Top-1 and neighbor-accuracy numbers.","section":"V-A, Tables I-II"},{"comment":"The proposed results are reported as the 'best result, after running our enhanced GA module 50 times on each image' (Section V-D), while the baseline methods are not described as receiving the same multiple-run treatment. With stochastic solvers, best-of-50 systematically inflates the expected reported accuracy relative to a single run, and the paper gives no error bars, confidence intervals, or significance tests. The very large gaps in Table II (e.g., 95.2% vs. 28% for Bridger et al.) could be partly an artifact of comparing a selected best run against a single run. Please report mean and standard deviation (or median and IQR) over the 50 runs for the proposed method, run each baseline for the same number of trials with the same stopping rule, and report paired significance tests (e.g., Wilcoxon signed-rank over images). The same issue applies to the erosion experiments in Table VI and the ablation results in Table III.","section":"V-C/D, Tables II, III, VI"},{"comment":"The paper does not specify how the 24 MNAz test panels are converted into 50x50 pieces for evaluation. In Section V-A, automated piece-cropping is described and immediately followed by the concession that 'automated cropping may not always align perfectly with actual piece boundaries.' If the test-set tiles are produced by the same automated procedure without manual verification, the ground-truth adjacency labels used to compute the neighbor accuracies in Table II may be incorrect, which would affect all compared methods to different degrees and undermine the SOTA comparison. The authors should state the exact cropping procedure for the test panels and, if automated, quantify the alignment error or provide manual verification for the 24 test panels.","section":"V-A, III-B"},{"comment":"For synthetic JPP, the manuscript reports 'average best results obtained over five runs of our scheme per image' (Section VI-A), while the Portuguese-tile section reports the best result over 50 runs; the aggregate statistic is ambiguous. Please clarify the number of runs and the aggregation rule used in Tables IV and V, and report variance so that the SOTA claims on synthetic puzzles can be assessed on the same footing as the baseline numbers.","section":"VI-A, Tables IV-V"}],"minor_comments":[{"comment":"Section V-D duplicates Section V-C almost verbatim (the text on elitism, roulette-wheel selection, phase-skip probabilities, and the 50-run protocol appears twice); the duplicate should be removed.","section":"V-D"},{"comment":"The sentence 'Although automated cropping may not always align perfectly with actual piece boundaries, such occurrences are rare and may contribute positively by reducing the risk of overfitting' is unconvincing: misaligned training crops reduce label quality rather than mitigating overfitting, and the claim should be reworded or supported.","section":"V-A"},{"comment":"Top-1 accuracies in Table I are reported as point estimates without any measure of variation across the 24 test panels; a per-panel standard deviation or confidence interval would strengthen the comparison with the baseline CMs.","section":"Table I"},{"comment":"The phrase 'our unique hybrid methodology' overstates novelty given that the authors' prior work [15] already combines a DL-based CM with a GA solver; the abstract should describe the specific extensions rather than claiming uniqueness.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The data-contamination risk is the most serious concern. The one-sentence statement that the test panels were 'excluded from CNN training' does not address near-duplicates in the 208 Internet images, and given that some images are casual tourist photos, this is not a purely theoretical concern. A deduplication check and clean-split re-evaluation should be a required part of the revision. The best-of-50 reporting is also a barrier to assessing the claimed advantage; the authors should be asked to provide mean/std or median results and comparable baseline runs. The duplicated Section V-D suggests the manuscript needs a careful editorial pass."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid engineering paper with a plausible hybrid DL+GA recipe and the widest set of experiments I've seen in one jigsaw-puzzle paper. The synthetic and eroded-boundary results look genuinely strong. The headline Portuguese-tile numbers, though, should not be taken at face value until the authors rule out train/test overlap.\n\nWhat's new and good: the DLCM ensemble (RGB, R, G, B sub-networks) is a sensible extension of the authors' prior GA+DL work, and the generic multi-domain framing is useful. They show competitive or SOTA neighbor accuracy on the standard MIT/McGill/Pomeranz synthetic benchmarks, and the eroded-boundary gains over Bridger et al. are large and consistent across 7% and 14% erosion. The 765-strip shredded-document reconstruction is a nice demo. The ablation study on GA phases is informative, and the fitness-deviation analysis is a good sanity check that the GA is near the objective's optimum.\n\nSoft spots, in increasing order of concern. First, minor: Sections V-C and V-D are near-duplicates, which looks like a copy-paste slip. Also minor: the paper says it 'curates and releases' a Portuguese tile benchmark, but I see no URL or availability statement; the release promise is unfulfilled in this version.\n\nSecond, comparison fairness: the authors report the best of 50 GA runs per image, while the baselines are presumably single deterministic runs. No error bars or significance tests are given. That doesn't invalidate the method, but it makes the enormous gaps on Portuguese tiles (95.2% vs 28% for Bridger) hard to interpret. The reader flagged this, and I agree.\n\nThird, and load-bearing: the train/test overlap question. The paper states the 24 MNAz test panels were excluded from CNN training, but it never says the 208 Internet-collected training images (some tourist photos) were deduplicated against the test set. If a test panel appears in a tourist photo, the DLCM can memorize specific piece-pair adjacencies rather than learn a generic compatibility function, directly inflating the Top-1 and neighbor-accuracy numbers. The paper's own acknowledgment that automated cropping 'may not always align perfectly' is a separate issue; the overlap risk is unaddressed. This needs to be resolved with a careful dedup check and, ideally, a released dataset so others can verify.\n\nOverall, the method is plausible, the synthetic and eroded results are credible, and the paper deserves a serious referee. But the SOTA claims on Portuguese tiles are only as strong as the held-out guarantee. I'd send it to peer review with a request for major revision: add the dedup analysis, report average and variance for all methods (not just the best-of-50 for the proposed), and make good on the dataset release.\n\nWho benefits: anyone working on jigsaw-puzzle reconstruction, heritage conservation, or document recovery. Worth reading group discussion, but with the overlap caveat on the table.","headline":"A capable hybrid DL+GA jigsaw solver with broad experiments; the Portuguese-tile SOTA claims are provisional until train/test overlap is ruled out.","tokens_in":20183,"tokens_out":2588,"would_cite":true,"duration_ms":27100,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Pairing a deep-learning compatibility model with a genetic solver reconstructs Portuguese tile panels at 95.2% and 89.4% neighbor accuracy, besting the prior eroded-puzzle method by up to 35.1 points.","keywords":["jigsaw puzzle problem","visual reconstruction","deep learning compatibility measure","genetic algorithm","Portuguese tile panels","eroded boundaries","shredded documents","convolutional neural networks"],"falsifier":"Re-score the 24 museum test panels using ground-truth tile boundaries obtained by human marking or high-resolution seam detection instead of the automated 50x50 crop, and recompute neighbor accuracy. If the 95.2% Type-1 and 89.4% Type-2 figures drop substantially, the reported state-of-the-art is an artifact of misaligned crop labels rather than a true reconstruction capability.","tokens_in":19212,"feed_emoji":"🧩","tokens_out":9687,"duration_ms":83839,"temperature":0.7,"pith_summary":"This paper argues that for real-world image reassembly tasks, the limiting factor is the compatibility measure: when tile edges are degraded or the imagery is repetitive, boundary-only color comparisons fail and greedy placement locks in early errors. The proposed remedy is a hybrid in which a compact convolutional network scores whole pairs of tiles as potential neighbors, and a genetic algorithm searches for the global arrangement using those scores. The authors report 95.2% and 89.4% neighbor accuracy on known-dimension Type-1 and Type-2 Portuguese tile panels, up to 96.9% on standard synthetic benchmarks, and average gains of 16.2 and 35.1 percentage points over the prior GAN-based method on 7% and 14% eroded puzzles. If this holds, the same two-part recipe—learned whole-piece compatibility plus evolutionary global search—applies across domains with little per-task engineering, and it could turn an archaeology-scale manual reassembly effort into a machine-assisted workflow.","feed_headline":"Neural plus genetic method reassembles 95% of tile panels","feed_subtitle":"Piece-pair CNN scores plus an evolutionary solver beat prior eroded-puzzle results by 16–35 percentage points.","key_machinery":"The load-bearing object is the DLCM–GA combination. The DLCM is a convolutional network that ingests a whole piece pair as a $P\\times 2P$ image rather than comparing boundary pixels, so the compatibility signal can come from interior texture, color, and structure; the GA then treats the resulting pairwise scores as a fitness landscape and searches globally with a hierarchical crossover that places tiles by parent agreement, best-buddy relations, and fallback compatibility, plus random phase-skipping as mutation. Post-processing of the score matrix—per-edge min–max normalization and symmetrization—is also load-bearing, since it lifts DLCM Top-1 accuracy from 64.5% to 69.9% on Type-1 panels.","core_discovery":"The paper's central claim is that a compatibility measure which sees entire pieces, not just their abutting edges, is enough to make large real-world jigsaw puzzles tractable. The proposed DLCM is a compact convolutional network that takes a concatenated pair of $P\\times 2P$ tiles and outputs a scalar score; for the tile domain it is an ensemble of four such networks, one per color channel plus an RGB network, trained with binary cross-entropy on sampled positive and negative pairs and augmented with boundary degradation and pixel shifts. Raw scores are min–max normalized per edge and symmetrized so $C(e_i,e_j)=C(e_j,e_i)$. The companion solver is a genetic algorithm whose crossover grows a kernel through hierarchical phases, including parent-confidence phases and a best-buddies phase, with mutation that skips phases to escape local optima. On the paper's evaluation this yields 95.2% and 89.4% known-dimension neighbor accuracy for Type-1 and Type-2 Portuguese tile panels, new best average results across the standard synthetic Type-1 and Type-2 benchmarks, and average gains of 16.2 and 35.1 percentage points over the previous GAN-based method on 7% and 14% eroded puzzles. The same pipeline reconstructs a 765-strip shredded-document puzzle at 97.1% accuracy.","pith_inferences":["Because the reconstruction accuracy (95.2% Type-1) far exceeds the DLCM's Top-1 compatibility accuracy (69.9% Type-1), the GA is not merely summing evidence—it is actively correcting many wrong first choices; a testable consequence is that improving the compatibility measure may matter less than improving the solver's global search in this regime.","The automated-cropping caveat applies to absolute accuracy; however, since every compared method is scored on the same labels, the relative gap over prior methods is more trustworthy than the headline numbers.","A direct stress test would evaluate on panels whose tile seams are known from the physical tiles or manually marked, so the sensitivity to crop alignment can be quantified.","The computational bottleneck the paper flags—computing $16N^2$ pairwise scores—means the practical ceiling on the number of pieces is set by the CNN, not the GA; embedding-based compatibility could extend the same recipe to tens of thousands of pieces."],"forward_implications":["Portuguese tile panels with known dimensions can be assembled automatically to near-perfect neighbor accuracy, reducing a decades-long manual effort to a machine-assisted task.","The same trained compatibility network transfers to synthetic jigsaw benchmarks and to eroded-boundary puzzles with only a change in training data, so the hybrid is a general recipe rather than a tile-specific method.","Heavily eroded pieces (14% of boundary pixels removed) remain reconstructible at 85–92% neighbor accuracy, roughly 35 points above the prior GAN-based method.","On strip-cut shredded documents, the framework reaches 97.1% accuracy on a 765-strip multi-page puzzle, enough to recover the text content.","The GA's stochastic restarts matter: best-of-50 accuracy is 95.2% Type-1 and 89.4% Type-2, while average-of-runs is 93.1% and 79%, so the reported state-of-the-art figures require multiple runs per puzzle."],"supporting_citations":[{"why":"Supplies the genetic-algorithm solver with hierarchical crossover that the paper adapts and ablates.","marker":"[13]"},{"why":"The authors' prior DL-plus-GA hybrid for Portuguese tile panels; this paper extends it and compares against its reported accuracies.","marker":"[15]"},{"why":"The GAN-based compatibility measure for eroded boundaries that is the main baseline on erosion experiments and on Portuguese panels.","marker":"[44]"},{"why":"Provides the Pomeranz benchmarks and the greedy solver used as a comparison method in the tile-panel and synthetic experiments.","marker":"[29]"},{"why":"Introduces the MGC compatibility measure and Type-2/unknown-dimension solving, a baseline the paper retrains and beats.","marker":"[30]"},{"why":"Provides the greedy solver with missing-piece handling that later methods build on and that the paper includes as a comparison.","marker":"[37]"},{"why":"Contributes the MIT dataset, one of the three standard synthetic benchmarks used for the Type-1 and Type-2 comparisons.","marker":"[27]"},{"why":"Supplies the McGill dataset used as a second synthetic benchmark in the comparisons.","marker":"[68]"},{"why":"Supplies the SqueezeNet-based deep learning compatibility approach that is retrained on Portuguese tiles and used in the CM comparison.","marker":"[40]"}],"fun_headline_variants":["Whole-tile CNN scores plus GA solver top real-world jigsaw benchmarks","Hybrid neural-genetic method reassembles Portuguese tiles at 95% accuracy","Pairwise neural compatibility guides evolutionary puzzle reconstruction","Eroded puzzles gain 16-35 points via hybrid CNN-GA framework","Deep compatibility model and genetic algorithm fix large degraded jigsaws"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim depends on the 24 Portuguese museum test panels being cut into tiles along their true boundaries: if the automatic 50x50 cropping is even slightly misaligned, the 'correct neighbor' labels are wrong, and the reported 95.2% and 89.4% accuracy would be inflated.","fun_headline_variants_meta":{"raw":{"variants":["Whole-tile CNN scores plus GA solver top real-world jigsaw benchmarks","Hybrid neural-genetic method reassembles Portuguese tiles at 95% accuracy","Pairwise neural compatibility guides evolutionary puzzle reconstruction","Eroded puzzles gain 16-35 points via hybrid CNN-GA framework","Deep compatibility model and genetic algorithm fix large degraded jigsaws"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000267,"raw_usage":{"total_tokens":1616,"prompt_tokens":946,"completion_tokens":670,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":562,"completion_tokens_details":{"reasoning_tokens":579}},"tokens_in":562,"tokens_out":670,"duration_ms":8073,"temperature":1.0,"reasoning_tokens":579,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T20:35:20.441297+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-score the 24 museum test panels using ground-truth tile boundaries obtained by human marking or high-resolution seam detection instead of the automated 50x50 crop, and recompute neighbor accuracy. If the 95.2% Type-1 and 89.4% Type-2 figures drop substantially, the reported state-of-the-art is an artifact of misaligned crop labels rather than a true reconstruction capability.","supporting_citations":[{"cited_title":"A genetic algorithm- based solver for very large jigsaw puzzles,","cited_arxiv_id":null,"evidence_quote":"Supplies the genetic-algorithm solver with hierarchical crossover that the paper adapts and ablates."},{"cited_title":"A novel hybrid scheme using genetic algorithms and deep learning for the reconstruction of Portuguese tile panels,","cited_arxiv_id":null,"evidence_quote":"The authors' prior DL-plus-GA hybrid for Portuguese tile panels; this paper extends it and compares against its reported accuracies."},{"cited_title":"Solving jigsaw puzzles with eroded boundaries,","cited_arxiv_id":null,"evidence_quote":"The GAN-based compatibility measure for eroded boundaries that is the main baseline on erosion experiments and on Portuguese panels."},{"cited_title":"A fully automated greedy square jigsaw puzzle solver,","cited_arxiv_id":null,"evidence_quote":"Provides the Pomeranz benchmarks and the greedy solver used as a comparison method in the tile-panel and synthetic experiments."},{"cited_title":"Jigsaw puzzles with pieces of unknown orientation,","cited_arxiv_id":null,"evidence_quote":"Introduces the MGC compatibility measure and Type-2/unknown-dimension solving, a baseline the paper retrains and beats."},{"cited_title":"Solving multiple square jigsaw puzzles with missing pieces,","cited_arxiv_id":null,"evidence_quote":"Provides the greedy solver with missing-piece handling that later methods build on and that the paper includes as a comparison."},{"cited_title":"A probabilistic image jigsaw puzzle solver,","cited_arxiv_id":null,"evidence_quote":"Contributes the MIT dataset, one of the three standard synthetic benchmarks used for the Type-1 and Type-2 comparisons."},{"cited_title":"A biologically inspired algorithm for the recovery of shading and reflectance images,","cited_arxiv_id":null,"evidence_quote":"Supplies the McGill dataset used as a second synthetic benchmark in the comparisons."},{"cited_title":"A deep learning-based compatibility score for reconstruction of strip-shredded text documents,","cited_arxiv_id":null,"evidence_quote":"Supplies the SqueezeNet-based deep learning compatibility approach that is retrained on Portuguese tiles and used in the CM comparison."}],"review_version":1}