{"id":"4cdf0e08-1dc6-4268-9825-a794eb5b653c","arxiv_id":"2411.14729","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A CNN-Transformer model with post-quantization compression is presented for detecting network attacks in cooperative smart farming, with claimed accuracies up to 97% on proprietary testbed data.","lead":"This paper proposes a CNN-Transformer network anomaly detector for edge devices in cooperative smart farming, testing it on two self-collected datasets. The authors also apply post-quantization compression and compare against traditional ML models, though the reported comparison contains an internal inconsistency.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own reported numbers contradict its central claim: in Smart Farm A the compressed CNN-Transformer has 90% accuracy versus CNN-LSTM's 92%, yet Section VII.D says it is 'higher'; the claimed superiority over baselines is not supported by the reported data.","rationale":"The reader's verdict is REJECT, and this stress-test pass agrees that a rejection is warranted. However, the most load-bearing concern is not primarily the representativeness of the testbed, but an internal contradiction in the paper's own comparison. The strongest claim is explicitly that the compressed CNN-Transformer outperforms all listed baselines in both accuracy and F1. The evidence presented for Smart Farm A is the sentence in Section VII.D: 'the accuracy of the compressed CNN-Transformer model is higher than that of the CNN-LSTM model, with 90% compared to 92%.' A model with 90% accuracy is not higher than one with 92%. Unless the figure, the sentence, or the underlying values are corrected, the central empirical claim is false as stated. No amount of external-validity repair can fix this: the numbers themselves do not support the conclusion. The absence of confidence intervals or repeated-seed runs makes it impossible to infer that the small reported differences on Smart Farm B are meaningful. This reinforces the reader's rejection without changing the verdict. The concrete test of re-deriving the comparison from raw outputs would settle whether the contradiction is a typo or a substantive error, but either way the paper as written cannot be accepted.","tokens_in":11801,"tokens_out":3882,"duration_ms":36262,"concrete_test":"Re-create the exact numeric values behind Figures 9 and 10 from the authors' prediction outputs (or, if unavailable, request the datasets and code). Compare compressed CNN-Transformer (5 encoders) accuracy and F1 against CNN-LSTM on the same Smart Farm A test split. If CNN-LSTM is 92% and CNN-Transformer is 90%, the claim in Section VII.D and Section VIII is false; if the values are reversed, verify which model the labels belong to and correct the text. This settles the concern directly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the compressed CNN-Transformer 'outperformed them [RF, LR, CNN, LSTM, CNN-LSTM] in both accuracy and F1 score' (Section VIII). The only direct evidence is Figures 9 and 10. Section VII.D states: 'For Smart Farm A, the accuracy of the compressed CNN-Transformer model is higher than that of the CNN-LSTM model, with 90% compared to 92%.' These numbers are inverted: 90% is lower than 92%, so the sentence as written is self-contradictory. On Smart Farm B the model reportedly beats CNN-LSTM by 3%, but no error bars, confidence intervals, or repeated-seed statistics are provided. Because the headline empirical claim depends on a comparison to strong baselines, one directly reported comparison refuting it is a load-bearing error: either the figure is mislabeled, the accuracy values are swapped, or the conclusion overstates the result. In all cases, the conclusion as stated is not justified by the evidence in the paper. This internal inconsistency is more decisive than the dataset-transferability concern: even if the testbed were perfectly representative, the reported numbers still do not show the claimed superiority.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a secure architecture for Cooperative Smart Farming (CSF) with edge-based digital twins, builds a two-farm testbed, collects two WiFi network datasets under normal and attack conditions, and develops a CNN-Transformer anomaly detector that is compressed with post-training quantization. The authors report accuracy and F1-score comparisons against Random Forest, Logistic Regression, CNN, LSTM, and CNN-LSTM, and claim that the compressed CNN-Transformer outperforms these baselines and that detection accuracy improves as the number of encoder layers increases.","tokens_in":11955,"tokens_out":4018,"duration_ms":38967,"significance":"If the empirical claims held, the paper would demonstrate a lightweight, edge-deployable network anomaly detector tailored to cooperative smart farming, with a concrete testbed and two new datasets. The testbed construction, the use of Azure Digital Twin at the edge, and the explicit reporting of model memory before and after compression are useful contributions. However, the central comparative claim is currently undermined by an internal contradiction in the reported accuracy numbers and by the absence of any uncertainty quantification, and the external validity of the private, self-generated datasets is not established. These issues must be resolved before the claimed superiority over baselines can be accepted.","major_comments":[{"comment":"The manuscript states in Section VII.D: 'For Smart Farm A, the accuracy of the compressed CNN-Transformer model is higher than that of the CNN-LSTM model, with 90% compared to 92%.' Since 90% is lower than 92%, this sentence contradicts itself, and the conclusion in Section VIII that the compressed CNN-Transformer 'outperformed them in both accuracy and F1 score' is not supported by the reported numbers as written. Please correct the values, the figure labels, or the conclusion, and provide the F1 scores for both farms in a table so the comparison can be verified.","section":"Section VII.D, Figs. 9-10, Section VIII"},{"comment":"All accuracy and F1 results are single-run point estimates presented without error bars, confidence intervals, or statistical significance tests. The claims that the compressed model 'outperforms' CNN-LSTM, that accuracy 'slightly improved' or 'gradually declined' with epochs, and that more encoder layers improve detection are not statistically grounded. Please run multiple trials with different random seeds and report means and standard deviations, or use appropriate significance tests, for the comparisons in Figures 6, 7, 9, and 10.","section":"Section VII.C and VII.D"},{"comment":"The digital twin attack is implemented by a Python script that generates 'simulated data vastly different from the physical counterpart,' but the manuscript does not show that this corruption is reflected in the WiFi network traffic features used by the CNN-Transformer model. If DT corruption does not alter the captured network traffic, the model cannot detect it. Please provide an analysis of how the DT attack changes the network-traffic features, or evaluate the model directly on DT data, to support the claim that the proposed approach detects digital twin attacks.","section":"Section V.B and V.C"},{"comment":"The conclusion that 'the model's ability to detect the network attack increases with the increasing number of encoder layers' is based on only 1, 3, and 5 encoder layers, and the reported results are not monotonic: for Smart Farm A with 3 encoder layers, accuracy 'slightly improved to 91% at 30 epochs but gradually declined as the number of epochs increased.' Please report complete results for all encoder-layer counts and epochs, with uncertainty, before making this general claim.","section":"Section VII.C and VIII"}],"minor_comments":[{"comment":"The sentence 'The training and evaluation are conducted on two edge devices using Google Colab' is contradictory, because Google Colab is a cloud-hosted notebook service, not an edge device. Please clarify which stages ran on the edge hardware and which ran in Colab.","section":"Section VII.B"},{"comment":"The dataset split into training, validation, and testing sets is mentioned but no split ratio is given. Please state the exact proportions used.","section":"Section V.C"},{"comment":"Reference [14] is cited for a 'robust intrusion detection system for DDoS attacks in smart agriculture,' but the listed title is about spatial prediction of soil organic carbon, which appears unrelated. Please verify the citation.","section":"Section II.B, reference [14]"},{"comment":"The phrase 'a normalized layer' should likely be 'a normalization layer' when describing the feed-forward block of the Transformer encoder.","section":"Section VI.A"},{"comment":"Reference [35] is titled 'AWS Managed Grafana' but the URL points to Azure Digital Twins documentation. Please correct the reference title and URL.","section":"Reference [35]"},{"comment":"The phrase 'various performance matrices' should be 'various performance metrics.'","section":"Section VII.C"}],"recommendation":"major_revision","confidential_remarks":"The decisive issue is the internal contradiction in Section VII.D, where the compressed CNN-Transformer is reported as 'higher' despite having 90% accuracy versus CNN-LSTM's 92%. If the actual measurements confirm the numbers as printed, the paper's central conclusion is false; if the numbers are typographical errors, the manuscript still needs repeated experiments and uncertainty quantification to support a superiority claim. I recommend major revision rather than immediate rejection because the flaw is local and potentially fixable, but the authors must provide corrected data and statistical support."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper builds a real testbed for cooperative smart farming, collects two new WiFi network datasets under a sensible set of attacks (deauth, SYN flood, ARP poisoning, DNS spoofing, evil twin, SQL injection, port scanning), and applies a standard CNN-Transformer with post-quantization for edge deployment. That is real work, and for a community drowning in pure simulation, the testbed and dataset collection are the most valuable part of the paper.\n\nThe problems start with the central empirical claim. In Section VII.D the authors write that the compressed CNN-Transformer 'is higher than that of the CNN-LSTM model, with 90% compared to 92%.' Those numbers say the opposite. Since the conclusion that the compressed model 'outperformed them all' rests on exactly this comparison, the conclusion is not supported by the reported data. No error bars, no repeated runs, no statistical tests anywhere else either; the figures plot single trajectories. The two datasets are not released, and no code is provided, so the main positive artifact—the datasets—cannot be used or checked.\n\nThe title and abstract promise detection of coordinated cyber and digital twin attacks, but the model only sees network traffic. The DT attack is simulated by injecting out-of-distribution values into the Azure DT, and no evidence shows those attacks alter the WiFi traffic that the model classifies. In fact, the future-work section says 'including DT datasets,' which reads as an admission that DT-specific data is not in the current evaluation.\n\nThere is also a small citation slip: reference [35] is titled 'AWS Managed Grafana' but points to the Azure Digital Twins overview. Minor, but symptomatic of the hurried write-up.\n\nWho gets value from this? Readers working on IoT/agricultural testbeds and attack implementation might mine the setup section. As a scientific claim about a lightweight detector, the paper is not currently defensible. The reported numbers contradict the headline result, and the empirical protocol is single-run and opaque. I would not cite it in its current form, and I would not use the numbers in a comparison. If the authors correct the comparison, release the datasets and code, and add uncertainty quantification, the testbed alone could support a useful paper. As it stands, it deserves a serious referee only because the testbed and dataset collection are genuine contributions that could be salvaged.","headline":"A useful testbed and two new smart-farming datasets, but the paper's central comparison is undercut by its own contradictory numbers, so the results as stated cannot be trusted.","tokens_in":12551,"tokens_out":2253,"would_cite":false,"duration_ms":21369,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A quantized CNN-Transformer model detects coordinated cyber and digital twin attacks in cooperative smart farming networks, reaching 97% accuracy while using roughly 90% less memory than the uncompressed model.","keywords":["cooperative smart farming","digital twin security","network anomaly detection","CNN-Transformer","post-training quantization","edge computing","IoT cyberattacks"],"falsifier":"Capture WiFi traffic while an attacker directly corrupts the Azure Digital Twin model state through its API rather than injecting out-of-distribution sensor values, and run the trained model on that traffic; if no anomaly is flagged, the assumption that DT corruption changes network traffic is false.","tokens_in":11523,"feed_emoji":"🌾","tokens_out":5117,"duration_ms":48456,"temperature":0.7,"pith_summary":"The paper claims that a lightweight, quantized CNN-Transformer can be deployed at the edge to detect network attacks on cooperative smart farms, including attacks that corrupt digital twins. The authors build a two-farm testbed, run eight types of network attacks across physical and digital twin layers, and capture two WiFi traffic datasets. On those datasets, the model with five transformer encoder layers reaches 97% accuracy, and post-training quantization cuts memory by about 90% with only a 1-2% accuracy drop. The central claim is that this compressed model outperforms traditional machine learning baselines while remaining small enough for edge devices.","feed_headline":"Compressed CNN-Transformer catches smart-farm attacks at 97%","feed_subtitle":"A quantized edge model shrinks memory about 90 percent while beating traditional machine-learning detectors.","key_machinery":"The central object is a hybrid CNN-Transformer architecture: a 1D CNN with batch normalization extracts local spatial patterns from network flow features, and a transformer encoder with multi-head self-attention and layer normalization captures long-range dependencies in the sequence. A global average pooling layer, a fully connected layer, and softmax perform the final attack classification. The load-bearing mechanism is the combination of local feature extraction by the CNN with the transformer's self-attention over distant time steps or features, which the paper argues is what makes the model effective at network anomaly detection. Post-training quantization to 8-bit integers is the mechanism that makes the model edge-deployable by shrinking memory usage roughly tenfold.","core_discovery":"On the paper's own terms, the core discovery is that a hybrid CNN-Transformer network anomaly detector, trained on WiFi traffic features from a cooperative smart farming testbed, can distinguish benign traffic from coordinated cyber and digital twin attacks with up to 97% accuracy. The model's detection ability improves as the number of transformer encoder layers increases from one to five, at the cost of greater memory use. Post-training quantization converts the weights from 32-bit floats to 8-bit integers, reducing memory from 4081.62 KB to 383.76 KB in the reported one-encoder case, while lowering accuracy by only about 1-2%. The compressed model is then shown to beat Random Forest, Logistic Regression, CNN, LSTM, and CNN-LSTM baselines in both accuracy and F1 score on the two self-generated datasets.","pith_inferences":["Beyond the paper: the detector only sees network traffic, so its ability to catch digital twin corruption depends on the attack changing WiFi traffic patterns; a DT attack that tampers with the virtual model internally without generating network chatter would be invisible to this approach.","Beyond the paper: training a single detector on traffic from many farms with different sensor configurations would require domain adaptation or federated learning, since the current model is trained per farm on farm-specific features.","Beyond the paper: the same architecture and quantization pipeline could be tested on public IoT or agricultural intrusion datasets to determine whether the 97% accuracy reflects the model's general capability or the specific testbed traffic."],"forward_implications":["If the accuracy results are correct, each farm's edge server can run the detector and isolate a compromised farm before corrupted sensor data reaches the cooperative cloud.","Increasing transformer encoder layers from 1 to 5 improves attack detection on both datasets, suggesting model depth is a useful lever even in edge settings once quantization offsets the memory cost.","The 1-2% accuracy loss under 8-bit quantization is small relative to the roughly 90% memory reduction, so the compressed model is the version suitable for edge deployment.","On the paper's two datasets, the compressed model outperforms RF, LR, CNN, LSTM, and CNN-LSTM in accuracy and F1 score, indicating the hybrid architecture retains an advantage after compression.","The two generated datasets and the eight attack types provide a reusable testbed for future cooperative smart farming security research."],"supporting_citations":[{"why":"Supplies the transformer-based intrusion detection approach that the CNN-Transformer model builds on.","marker":"[17]"},{"why":"Provides a robust transformer-based intrusion detection system used as motivation for self-attention in network anomaly detection.","marker":"[18]"},{"why":"Supplies the rationale and method for edge deployment through quantization and knowledge distillation.","marker":"[19]"},{"why":"Provides the post-quantization training technique applied to IoT intrusion detection, which the paper adapts for its model.","marker":"[21]"},{"why":"Supplies quantization and compression background for transformer models, justifying the memory-reduction strategy.","marker":"[22]"},{"why":"Prior work on hierarchical federated transfer learning and digital twin security in cooperative smart farming that this paper extends and compares against.","marker":"[26]"},{"why":"Provides deep learning intrusion detection baselines in agriculture 4.0 that inform the comparison with traditional ML models.","marker":"[27]"}],"fun_headline_variants":["Edge AI spots coordinated farm attacks at 97%","Quantized transformer thwarts smart-farm cyberattacks","97% accuracy with 90% memory cut for edge defense","CNN-Transformer beats ML on farm cyberattack detection","Tiny edge model detects cyber and twin farm attacks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a compromised digital twin's corruption shows up in the WiFi network traffic captured on the testbed; if a real digital-twin attack leaves network traffic unchanged, the detector cannot see it.","fun_headline_variants_meta":{"raw":{"variants":["Edge AI spots coordinated farm attacks at 97%","Quantized transformer thwarts smart-farm cyberattacks","97% accuracy with 90% memory cut for edge defense","CNN-Transformer beats ML on farm cyberattack detection","Tiny edge model detects cyber and twin farm attacks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000276,"raw_usage":{"total_tokens":1662,"prompt_tokens":979,"completion_tokens":683,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":604}},"tokens_in":595,"tokens_out":683,"duration_ms":30344,"temperature":1.0,"reasoning_tokens":604,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:57:56.942458+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Capture WiFi traffic while an attacker directly corrupts the Azure Digital Twin model state through its API rather than injecting out-of-distribution sensor values, and run the trained model on that traffic; if no anomaly is flagged, the assumption that DT corruption changes network traffic is false.","supporting_citations":[{"cited_title":"Intrusion detection model based on improved transformer,","cited_arxiv_id":null,"evidence_quote":"Supplies the transformer-based intrusion detection approach that the CNN-Transformer model builds on."},{"cited_title":"Rtids: A robust transformer- based approach for intrusion detection system,","cited_arxiv_id":null,"evidence_quote":"Provides a robust transformer-based intrusion detection system used as motivation for self-attention in network anomaly detection."},{"cited_title":"Qae-ids: Ddos anomaly detection in iot devices using post-quantization training,","cited_arxiv_id":null,"evidence_quote":"Provides the post-quantization training technique applied to IoT intrusion detection, which the paper adapts for its model."},{"cited_title":"Hierarchical federated transfer learning and digital twin enhanced secure cooperative smart farming,","cited_arxiv_id":null,"evidence_quote":"Prior work on hierarchical federated transfer learning and digital twin security in cooperative smart farming that this paper extends and compares against."},{"cited_title":"Deep learning- based intrusion detection for distributed denial of service attack in agriculture 4.0,","cited_arxiv_id":null,"evidence_quote":"Provides deep learning intrusion detection baselines in agriculture 4.0 that inform the comparison with traditional ML models."}],"review_version":1}