{"id":"d1b54e01-705d-4713-9a4d-b333601e7325","arxiv_id":"2412.06643","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"Extending MesoNet to six convolutional layers and swapping the deepfake training images for Stable Diffusion images raises three-class accuracy from 43% to 76% on a small in-distribution dataset, without external validation.","lead":"This paper trains several small neural networks to tell real face photos apart from AI-generated or face-swapped images, and reports up to 76% accuracy on a small custom dataset. It is an incremental engineering study built on the existing MesoNet architecture, not a breakthrough in deepfake detection.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"76% accuracy is not shown to be out-of-sample: no train/test split is documented and the DeepFake class was replaced by Stable Diffusion images, so the reported result does not support the claimed superiority.","rationale":"The reader's REJECT verdict is supported. The reported 76% is internally consistent, but the paper gives no evidence that the model's predictions transfer to images not used in training. The absence of a documented data split is the sharpest technical form of this problem: Section IV-A says the balanced database is used \"to train and evaluate\" the multi-classifier, and the support counts in Table VIII match the full dataset sizes, which is consistent with evaluating on training data. In addition, replacing the DeepFake class with Stable Diffusion images (Section IV-A, Table V) changes the semantic content of the class while the label \"DeepFake\" is retained in Table VIII. The improvement from 68% to 76% may therefore reflect separation of DiffusionDB-specific artifacts rather than detection of facial manipulations. The paper does report honest baselines for its own earlier models (43%, 68%), and the architectural progression is coherent, but neither the 76% number nor the claimed superiority over traditional approaches is externally validated. A single held-out replication would settle whether the central claim survives. This does not change the reader's verdict; it reinforces it.","tokens_in":11539,"tokens_out":4004,"duration_ms":44375,"concrete_test":"Obtain or reconstruct the exact train/validation/test partition used for Table VIII; if none exists, re-run the MesoMultiNet+ pipeline with a strict, stratified 80/20 train/test split over the 4,480 images in Table V and report test accuracy on the held-out 20%. If held-out accuracy is materially below 76%, the reported result is in-sample and the headline claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—76% accuracy in Section IV-F/Table VIII—rests on an evaluation that never demonstrates out-of-sample performance. Section IV-A presents a single \"balanced dataset\" used \"to train and evaluate\" the multi-class system and gives no train/validation/test split; Section IV-F reports metrics on the same modified dataset after the DeepFake class was replaced by Stable Diffusion images from DiffusionDB (Table V). With no held-out partition, the 76% value may be in-sample memorization rather than discrimination of manipulations. Even if a split exists, the replacement changes the task: the \"DeepFake\" row in Table VIII is actually Stable Diffusion images, so the result cannot support the abstract's claim that the model \"surpasses traditional approaches\" on manipulated facial images—no external method is compared and no unseen-manipulation test is run. This is the load-bearing weakness: the one quantitative result that would justify the paper does not establish what it claims.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes lightweight CNN architectures for detecting manipulated facial images, extending the MesoNet approach. The authors introduce a binary model (MesoNet+) and two multi-class models (MesoMultiNet and MesoMultiNet+) with four and six convolutional layers, respectively. They evaluate the binary model on the original MesoNet dataset and Celeb-DF, and the multi-class models on a balanced dataset of bonafide, DeepFake (later replaced by Stable Diffusion images from DiffusionDB), and FaceSwap images. The headline result is 76% accuracy for the six-layer model on the modified dataset, compared to 43% and 68% for the two earlier multi-class variants. The paper also documents the original MesoNet's generalization failure on Celeb-DF (AUC drop from 0.96 to 0.56).","tokens_in":11773,"tokens_out":6272,"duration_ms":56971,"significance":"If the 76% accuracy were established under a valid evaluation protocol, the proposed models would offer a practically useful lightweight solution for distinguishing real faces from certain AI-generated manipulations and face swaps, with potential deployment in resource-constrained settings. The paper's strength lies in its transparent reporting of negative results: the drop in MesoNet's AUC on Celeb-DF and the initial 43% accuracy of the four-layer multi-class model are honestly presented. The explicit disclosure of the dataset modification (replacing DeepFake images with Stable Diffusion images) is also a point in the authors' favor. However, the central quantitative claim is not currently supported: the final experiment changes the classification task, no held-out test set is documented, and no external baseline is compared, so the reported 76% does not substantiate the abstract's claim of surpassing traditional approaches.","major_comments":[{"comment":"In Section IV-A you state that, because FaceSwap and DeepFake images were difficult to differentiate, the DeepFake images from the previous set were replaced by images generated with Stable Diffusion from the DiffusionDB dataset. Consequently, the 'DeepFake' class in Tables V, VII, VIII, and IX is not a deepfake class at all; it is a Stable Diffusion class. The 76% accuracy reported in Section IV-F and Table VIII is therefore for a three-class problem of bonafide vs. FaceSwap vs. Stable Diffusion images, not for the task claimed in the abstract (distinguishing manipulated images from genuine ones). The 'DeepFake' row in Table VIII is mislabeled, and the improvement from 68% to 76% in Table IX is partly a result of this dataset change rather than of the architectural changes alone. Please either relabel the class and qualify the claims, or rerun the experiment with actual DeepFake images, or provide a clear justification for why Stable Diffusion images are a proxy for deepfakes.","section":"Section IV-A, Tables IV-V and Section IV-F, Table VIII"},{"comment":"No train/validation/test split is described for the multi-classification experiments. Section IV-A says the balanced dataset is used to train and evaluate the multi-classifier system, and Table V lists only total image counts per class; Section IV-F then reports the 76% accuracy apparently on the same data. With roughly 1,500 images per class and a six-layer CNN, the reported metrics may reflect in-sample performance rather than generalization to unseen images. Please specify the exact data partitioning (e.g., a fixed train/validation/test split or cross-validation), state which partition produced the numbers in Tables VI-VIII, and verify that no test images were used for model selection or early stopping. Without this, the 76% figure cannot be interpreted as an out-of-sample result.","section":"Section IV-A and Section IV-F"},{"comment":"The abstract claims the proposed models achieve accuracy surpassing traditional approaches, but Table IX compares only the three proposed architectures (MesoMultiNet, MesoMultiNet+, and MesoMultiNet+ with Stable Diffusion). No external baseline is evaluated under the same protocol: the original MesoNet (or MesoInception-4) is not retrained or evaluated on this dataset, and no published state-of-the-art forensics method is included. The progression 43% to 68% to 76% also conflates architectural changes with dataset changes, since the dataset composition changes between Table IV and Table V. A fair comparative evaluation against at least the original MesoNet on the identical data split is needed to support the superiority claim.","section":"Section IV-G and abstract"}],"minor_comments":[{"comment":"The ROC curve is defined with the false negative rate FPR; this should be the false positive rate.","section":"Section IV-B"},{"comment":"The phrase 'can be seen in Table I' likely refers to Table VI (or VII) for multi-class metrics; Table I is the binary dataset split.","section":"Section IV-D, final paragraph"},{"comment":"The model is named inconsistently: MesoMultiNet and MultiMesoNet are used interchangeably (e.g., Section IV-G vs. Section III-B); please standardize.","section":"Throughout"},{"comment":"Reference [11] is cited for MesoNet architecture details, but [11] is a ControlFace paper; the correct reference for MesoNet is [6]. The reference numbering appears misaligned.","section":"Section II-D"},{"comment":"The text says 'The original accuracy, which stood at 0.96, decreases drastically to 0.56' and then refers to an AUC-ROC value of 0.56; these are two different metrics and should not be conflated.","section":"Section IV-B"},{"comment":"There are several typos: 'DifussionDB' (Section IV-A), 'deconvolution layers' (Sections IV-B/IV-C; the added layers are convolutional), and 'We would like to thanks' (Acknowledgments).","section":"Various"}],"recommendation":"major_revision","confidential_remarks":"The paper is an incremental engineering study; the main result is compromised by the post-hoc substitution of the DeepFake class. I would ask the authors for a rigorous held-out evaluation and a direct comparison with the original MesoNet before considering publication. The manuscript may be better suited to a workshop or technical report venue in its current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take on arXiv:2412.06643: it is a routine MesoNet extension with an unsupported headline number. The 76% accuracy in the abstract and Section IV-F is measured on a self-made balanced dataset where the DeepFake class was replaced by Stable Diffusion images after the authors found FaceSwap vs DeepFake hard to distinguish. No train/test split is documented anywhere, and the paper never compares against any external detection method.\n\nWhat is actually there: the authors add two conv layers to MesoNet (MesoNet+), then adapt the architecture to multi-class (MesoMultiNet, MesoMultiNet+). The progression from 43% to 68% to 76% is internally consistent, and the paper honestly reports that original MesoNet collapses from 0.96 AUC on its training distribution to 0.56 on Celeb-DF. That is a real, reproducible observation about generalization.\n\nThe soft spots are not minor. The central claim rests on an evaluation that never demonstrates out-of-sample performance. Section IV-A describes a single balanced dataset \"to train and evaluate\" without specifying a split. Section IV-F reports metrics on the same dataset after the class swap, so the numbers could reflect in-sample memorization. The class swap itself is a post-hoc change to the task: the \"DeepFake\" row in Table VIII is actually Stable Diffusion images, so the result says nothing about detecting DeepFakes. The abstract's \"surpassing traditional approaches\" is unsupported because the only baselines are the paper's own earlier variants.\n\nThe paper ships no code, no data, and no error bars. The writing has rough edges (Spanish residual words, \"completeness\" for recall), but the main issue is evaluation validity, not style.\n\nWho this is for: someone tracking lightweight CNN variants for low-resource deepfake screening might skim it, but I would not rely on the numbers. It could serve as a cautionary teaching example of dataset-adaptive evaluation, but it does not deserve a serious referee as a research contribution.\n\nMy recommendation: desk reject. The authors need a proper held-out split, an external benchmark (e.g., FaceForensics++ or Celeb-DF), and a comparison to at least one existing detector before the 76% claim can be taken seriously.","headline":"Routine MesoNet extension whose headline 76% accuracy is undermined by a post-hoc dataset swap and no documented held-out split.","tokens_in":12264,"tokens_out":2264,"would_cite":false,"duration_ms":22878,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a six-layer convolutional network, MesoMultiNet+, distinguishes genuine faces from deepfake and face-swap manipulations with 76% accuracy on a balanced dataset, and that the improvement comes from deeper convolution…","keywords":["deepfake detection","facial image manipulation","convolutional neural network","MesoNet","multi-class classification","stable diffusion images","face swap detection","transfer learning"],"falsifier":"Run the trained MesoMultiNet+ model on an unseen deepfake dataset, such as the original DeepFake images from the MesoNet dataset or an independent benchmark, without any retraining; if accuracy falls toward chance levels (50% in binary or 33% in three-class settings), the reported 76% reflects dataset-specific artifacts rather than genuine generalization.","tokens_in":11380,"feed_emoji":"🕵️","tokens_out":4117,"duration_ms":38252,"temperature":0.7,"pith_summary":"The paper sets out to build lightweight convolutional neural networks that can tell genuine facial images apart from manipulated ones, extending the compact MesoNet architecture to handle multiple attack types at low computational cost. It reports that a six-layer multi-class model, MesoMultiNet+, reaches 76% accuracy on a balanced dataset separating real faces from deepfake and face-swap images, compared with 43% for the four-layer version and 68% for the six-layer version trained without AI-generated images. The authors argue that adding convolutional depth, variable filter sizes, facial-landmark alignment, and transfer learning from a pretrained binary model are what make this performance possible. A sympathetic reader would take the central claim to be that a modestly deeper CNN, trained with synthetic text-to-image outputs in the mix, can serve as an effective and resource-efficient detector of facial manipulation.","feed_headline":"A six-layer CNN tells real faces from fakes at 76 percent accuracy","feed_subtitle":"Lightweight MesoNet-based model separates genuine images from deepfakes and face swaps on a balanced test set.","key_machinery":"The load-bearing object is MesoMultiNet+, a six-layer convolutional architecture inspired by MesoNet. It uses alternating convolutional, batch-normalization, and max-pooling blocks with increasing filter counts and mixed kernel sizes, followed by dense layers with dropout and a softmax output. Training relies on categorical cross-entropy with the Adam optimizer, transfer learning from a pretrained binary MesoNet model, and a preprocessing step that aligns faces using a 68-point facial landmark detector before classification.","core_discovery":"The central discovery claimed in the paper is that extending the MesoNet architecture from four to six convolutional layers, together with replacing the deepfake training class with Stable Diffusion images, lifts multi-class facial manipulation detection from 43% to 76% accuracy on a balanced three-way dataset. In the final configuration, MesoMultiNet+ achieves per-class F1 scores of 0.78 for deepfake and face-swap and 0.74 for genuine images, with precision and recall balanced across all three classes. The authors also report that the binary MesoNet+ model reaches 90% accuracy on a balanced dataset of deepfakes and genuine images, and they attribute this improvement to the added layers and to careful preprocessing. The paper frames this as surpassing traditional approaches, where the comparison baseline is the original MesoNet binary model and its four-layer multi-class adaptation.","pith_inferences":["The headline 76% figure does not measure detection of classic deepfakes: the deepfake class in that experiment consists of Stable Diffusion images, so the number reflects separation among genuine, face-swap, and diffusion-generated faces rather than detection of the original DeepFake artifacts.","A natural next experiment, not reported in the paper, would be to evaluate the same six-layer architecture on an independent deepfake benchmark without retraining; that would show whether the accuracy transfers beyond the specific balanced dataset.","The claim of surpassing traditional approaches is supported only against the paper's own MesoNet baselines; a comparison with other post-MesoNet detection methods is absent, so the relative standing of the model in the broader field is not yet established.","The binary 90% accuracy on the balanced dataset is also dataset-specific; real-world value would depend on the prevalence of manipulations and the variety of unseen attack types encountered in deployment."],"forward_implications":["A compact six-layer CNN can separate genuine faces from deepfake and face-swap images at 76% accuracy on a balanced set, making multi-class manipulation screening feasible on low-resource devices.","The jump from 43% accuracy with four layers to 68% with six layers suggests that convolutional depth adds useful discriminative power for fine-grained manipulation artifacts.","Including Stable Diffusion images in the training set raised accuracy from 68% to 76%, indicating that exposure to text-to-image synthetic outputs sharpens the model's ability to separate manipulated from genuine faces.","Transfer learning from a pretrained binary model to the multi-class model accelerates training and improves feature reuse, according to the paper's reported results."],"supporting_citations":[{"why":"Provides the original MesoNet compact architecture that this work extends from four to six convolutional layers.","marker":"[6]"},{"why":"Celeb-DF supplies genuine and deepfake images used for the unknown-dataset evaluation and for the multi-class training and test sets.","marker":"[15]"},{"why":"DiffusionDB provides the Stable Diffusion images that replace the DeepFake class in the final 76% accuracy experiment.","marker":"[17]"},{"why":"The DFDC FaceSwap dataset supplies the FaceSwap attack class used in the multi-class classification datasets.","marker":"[16]"},{"why":"The 68-point facial landmark detector is used in the preprocessing pipeline to align faces before feeding them to the network.","marker":"[14]"}],"fun_headline_variants":["Six-layer CNN tells real from fake faces at 76%","MesoNet with extra layers catches 76% of facial edits","Deepfake detector: six-layer CNN beats baseline by 33%","CNN upgrade lifts face manipulation detection to 76%","Binary model hits 90% on deepfakes vs genuine"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results are computed on the same balanced dataset used for training, where the deepfake class was replaced with Stable Diffusion images; the assumption is that accuracy on this dataset reflects the model's ability to detect manipulations in the wild.","fun_headline_variants_meta":{"raw":{"variants":["Six-layer CNN tells real from fake faces at 76%","MesoNet with extra layers catches 76% of facial edits","Deepfake detector: six-layer CNN beats baseline by 33%","CNN upgrade lifts face manipulation detection to 76%","Binary model hits 90% on deepfakes vs genuine"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000588,"raw_usage":{"total_tokens":2732,"prompt_tokens":887,"completion_tokens":1845,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":503,"completion_tokens_details":{"reasoning_tokens":1759}},"tokens_in":503,"tokens_out":1845,"duration_ms":15427,"temperature":1.0,"reasoning_tokens":1759,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:26:09.577634+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained MesoMultiNet+ model on an unseen deepfake dataset, such as the original DeepFake images from the MesoNet dataset or an independent benchmark, without any retraining; if accuracy falls toward chance levels (50% in binary or 33% in three-class settings), the reported 76% reflects dataset-specific artifacts rather than genuine generalization.","supporting_citations":[{"cited_title":"Shape predictor 68-point facial landmark detector,","cited_arxiv_id":null,"evidence_quote":"The 68-point facial landmark detector is used in the preprocessing pipeline to align faces before feeding them to the network."}],"review_version":1}