{"id":"2ade152c-340f-4f3e-9f22-251792cf4c34","arxiv_id":"2508.08339","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":2.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The submitted body is an unrelated survey, not the SHeRL-FL method claimed in the metadata.","lead":"This preprint is internally inconsistent: the abstract describes a federated split-learning method, SHeRL-FL, but the full text is a survey on TinyML object detection. The claimed communication reductions in the abstract are not present in the submitted body.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim of SHeRL-FL's communication-efficiency gains is unsupported: the full text is a different TinyML survey with no mention of SHeRL-FL or its experiments.","rationale":"The reader's verdict (UNVERDICTED) is driven by the structural mismatch between the abstract and the full text. My stress-test converges on the same load-bearing concern: the abstract's headline result is a quantitative communication-efficiency claim, but the manuscript body contains neither the SHeRL-FL algorithm nor the experiments that would substantiate it. This is the single most important issue because, if the abstract is taken at face value, the paper promises a concrete algorithmic contribution with empirical validation; the full text instead delivers an unrelated survey. Without the method and experiments, there is no way to check the 90% and 50% reductions against baselines, datasets, or training configurations. The concern is not about consensus or novelty; it is about the presence of the evidence needed for the central claim to be meaningful. I do not see a way to rehabilitate the abstract's claim from the submitted body. However, I also do not think the correct conclusion is REJECT in the sense of 'the claim is false'—it is more accurately UNVERDICTED, because no assessment can be made. Since the reader already assigned UNVERDICTED, my read does not change the verdict; it reinforces it. A single concrete search/verification step would settle the matter: if the body truly contains no SHeRL-FL content, the abstract's central claim is unsupported by the submission as it stands.","tokens_in":45080,"tokens_out":2610,"duration_ms":33006,"concrete_test":"Search the full-text source (LaTeX/PDF) for the strings 'SHeRL-FL', 'SplitFed', 'HierFL', 'HAM10000', 'ISIC-2018', and 'CIFAR'. If 'SHeRL-FL' and 'SplitFed' are absent from the body and no experimental section reports the claimed datasets or communication measurements, the abstract's central claim is unsupported by the submission. Additionally, verify the arXiv metadata/PDF association: if the uploaded PDF is the TinyML survey, the mismatch is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central quantitative claim—that SHeRL-FL reduces data transmission by over 90% versus centralized FL and HierFL, and by 50% versus SplitFed, with experiments on CIFAR-10/100, HAM10000, and ISIC-2018—has no corresponding evidence in the submitted full text. The full text is a TinyML object-detection survey, titled 'Designing Object Detection Models for TinyML', with different authors and no occurrence of SHeRL-FL, split learning, hierarchical FL, or the claimed datasets. Thus the condition required for the abstract's claim to be assessable—that the manuscript contains the proposed method and its experiments—fails. This is not a subtle modeling assumption; the supporting artifact is absent. Even if the survey body is internally accurate and useful, it does not provide the algorithmic description, training setup, baselines, or measured transmission numbers needed to evaluate the claimed 90%/50% reductions. The submission therefore cannot support its own abstract, and any verdict on the central claim must remain unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission, as indexed, claims to propose SHeRL-FL, a method integrating split learning and hierarchical federated learning with representation learning, and reports communication reductions of over 90% versus centralized FL/HierFL and 50% versus SplitFed, with experiments on CIFAR-10, CIFAR-100, HAM10000, and ISIC-2018. The full text, however, is not that paper. It is a survey titled \"Designing Object Detection Models for TinyML: Foundations, Comparative Analysis, Challenges, and Emerging Solutions\" by different authors, covering quantization, pruning, knowledge distillation, and neural architecture search for object detection on microcontrollers. The body contains no mention of SHeRL-FL, split learning, hierarchical FL, or any of the claimed datasets. The central contribution and experimental evidence promised in the abstract are therefore absent from the manuscript.","tokens_in":45338,"tokens_out":3009,"duration_ms":35525,"significance":"If the abstract's claims were supported by a real method and experiments, SHeRL-FL would be a potentially significant communication-efficiency contribution to federated and split learning. However, as submitted, no such method, derivation, or experiment exists in the manuscript, so the significance of the claimed contribution cannot be assessed. Taken on its own terms as a TinyML object-detection survey, the full text has some merits: it offers a structured taxonomy of four optimization families, tabulated model comparisons (Tables 5, 6, 8), and a public repository link. These are useful survey elements, but they do not constitute the federated-learning paper described by the abstract.","major_comments":[{"comment":"The abstract promises a method called SHeRL-FL with quantitative results: 'reduces data transmission by over 90% compared to centralized FL and HierFL, and by 50% compared to SplitFed', based on experiments on CIFAR-10, CIFAR-100, HAM10000, and ISIC-2018. The full text is a different paper with a different title and author list: a TinyML object-detection survey. I could not find the term 'SHeRL-FL' anywhere in the body, nor any discussion of split learning, hierarchical federated learning, or the claimed datasets. The central claim of the submission is therefore unsupported by any content in the manuscript. This is a load-bearing mismatch: it is not a local error but the absence of the paper itself.","section":"Abstract; Sections 1 and 9"},{"comment":"There is no algorithmic description, training setup, aggregation rule, loss function, baseline configuration, or measured transmission-cost analysis for SHeRL-FL. The body instead surveys quantization, pruning, knowledge distillation, and NAS for object detection. Because the method and experiments are absent, the abstract's assertions about reduced coordination complexity and communication overhead cannot be derived, reproduced, or checked. This is not a gap that can be fixed by adding a missing section; it requires the actual SHeRL-FL manuscript.","section":"Sections 2-9"},{"comment":"Even if the TinyML survey is considered the intended submission, the stated research gap is disjoint from the abstract's claimed contribution. Section 1 defines the gap as the lack of surveys covering optimization techniques for OD on resource-constrained devices, while the abstract frames the contribution as a new FL/SL method. A paper cannot simultaneously be a novel federated-learning algorithm and a survey of TinyML object-detection compression without any connection between the two. The title, abstract, and full text describe different papers.","section":"Section 1"}],"minor_comments":[{"comment":"The text contains several typographical errors, e.g., '150,55 billion' (should be '150.55 billion'), 'compartive', and later 'qantization' and 'improvment'. These should be corrected in any revision.","section":"Section 1"},{"comment":"The taxonomy figure is dense and the subcategories are not numbered in the figure, making it hard to map to the text. Also, the self-citation [164] in Section 6.4.4 is used as an illustrative HNAS example; this is acceptable, but the authors should ensure it is clearly positioned as an example, not as a substitute for a broader literature discussion.","section":"Fig. 3; Section 6.4.4"}],"recommendation":"reject","confidential_remarks":"The mismatch between the metadata/abstract and the full text is a submission-integrity issue. The editor should verify whether the wrong PDF was uploaded and, if appropriate, contact the authors. Even under the most charitable reading, the manuscript does not support its own central claim, so it cannot be accepted in its current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The abstract and the full text are two different papers. The abstract describes SHeRL-FL, a split hierarchical FL method with claimed communication reductions; the body is a TinyML object-detection survey by different authors. You cannot evaluate the claimed result because the supporting artifact is absent.\n\nTaking the body on its own, it is a competent survey. It organizes four optimization families (quantization, pruning, KD, NAS) for deploying OD models on microcontrollers, includes comparison tables (Tables 5, 6, 8), a taxonomy figure, and a pointer to a public repository. That is honest incremental synthesis, useful to practitioners entering TinyML OD. The self-citation to the authors' own HNAS framework is an example, not load-bearing.\n\nThe mismatch between abstract and body is a load-bearing flaw, not formatting. The survey's tabulated benchmark numbers are transcribed from cited papers but not verified here; there are typos like '150,55 billion' and some missing math symbols. If the survey is to be used as a reference, those need cleanup. But the main issue is that the submission cannot support its own abstract; the 'SHeRL-FL reduces data transmission by over 90%...' claim has zero supporting description, experiments, or equations in the manuscript.\n\nThe survey might be a reasonable resource for people working on TinyML object detection, but as submitted this is not a coherent paper. A serious editor should desk reject; the authors need to resubmit the right full text or withdraw the abstract's claims.","headline":"The abstract and full text are two different papers—SHeRL-FL's claimed results are entirely absent from a competent but unrelated TinyML OD survey.","tokens_in":45820,"tokens_out":1902,"would_cite":false,"duration_ms":21991,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The manuscript pairs a SHeRL-FL abstract with a TinyML survey body.","keywords":["TinyML","object detection","model compression","quantization","pruning","knowledge distillation","neural architecture search","split learning"],"falsifier":"A reader can settle the mismatch immediately by searching the full text for 'SHeRL-FL'; it appears nowhere in the body, and no algorithm, no CIFAR/HAM10000/ISIC experiment, and no transmission-volume table is reported, so there is nothing in the manuscript that could confirm or refute the abstract's numbers.","tokens_in":44991,"feed_emoji":"📄","tokens_out":15925,"duration_ms":155280,"temperature":0.7,"pith_summary":"The submission as received has two mismatched layers. The metadata and abstract claim a new framework, SHeRL-FL, that combines split learning and hierarchical federated learning with representation learning at intermediate layers, and report that it cuts data transmission by more than 90% versus centralized FL and HierFL and by 50% versus SplitFed, with experiments on CIFAR-10/100, HAM10000, and ISIC-2018. The full text, however, is a different paper: a survey titled 'Designing Object Detection Models for TinyML' that analyzes quantization, pruning, knowledge distillation, and neural architecture search for object detection on microcontrollers, with comparison tables of MCU-deployed detectors. Read sympathetically, the full-text authors are trying to establish that existing surveys overlook optimization challenges specific to deploying object detectors on TinyML devices, and that a structured map of these techniques, plus benchmark numbers, fills that gap. If the abstract's SHeRL-FL claim were true, it would be a substantial communication-efficiency gain for hierarchical split learning; but the body contains no SHeRL-FL algorithm, no such experiments, and no transmission measurements.","feed_headline":"Promised split-learning paper contains a TinyML survey instead","feed_subtitle":"The full text never mentions SHeRL-FL, so the claimed 90% and 50% communication savings rest on missing experiments.","key_machinery":"For the abstract's claimed framework, the key mechanism is representation learning at intermediate layers: clients and edge servers would compute training objectives independently of the cloud, reducing coordination complexity and cutting the volume of data crossing tiers. For the full-text survey, the carrying mechanism is a taxonomy that splits model optimization into parameter removal (pruning), parameter quantization (QAT/PTQ/BNNs), parameter search (NAS), and knowledge transfer (KD), all applied to the backbone–neck–head pipeline of object detectors. The load-bearing evidence is Table 8, which compares MCU-deployed detectors by parameters, MMACs, peak SRAM, and mAP on PASCAL VOC; the ta","core_discovery":"The full-text authors' central claim is that previous surveys of lightweight object detection focus on backbones or general edge AI and miss the optimization of detection models under TinyML memory budgets. To close that gap, the survey organizes the field into four compression families—quantization (QAT, PTQ, binary networks), pruning (unstructured, structured, semi-structured), knowledge distillation (feature, multi-teacher, multi-modal, self-, weakly supervised), and neural architecture search (RL-, evolutionary-, gradient-, and hardware-aware)—and compares MCU-optimized detectors on PASCAL VOC, reaching 51.4–74.9% mAP with 53–511 kB peak SRAM. The body does not contain the SHeRL-FL frame","pith_inferences":["The claimed 50% cut versus SplitFed suggests the bottleneck it addresses is not body computation but the cross-tier transfer of intermediate activations; a natural follow-up experiment would measure per-round transmission as the cut layer moves through the network.","The survey's benchmark numbers come from different papers with different training setups, so a fair comparison of MCU detectors would require re-running the same models under a common training and quantization pipeline.","The taxonomy's emphasis on hardware-aware NAS and co-design suggests that future TinyML object detectors will increasingly be optimized jointly with the inference engine and memory scheduler, not just the network weights.","Treating the submitted file as two separate documents—a SHeRL-FL systems paper and a TinyML survey—would let each be evaluated on its own evidence."],"forward_implications":["The survey's taxonomy gives TinyML practitioners a direct way to match a compression family (e.g., quantization or NAS) to a specific hardware constraint such as peak SRAM or MMACs.","The tabulated MCU detectors show a current operating band of roughly 51–75% mAP on PASCAL VOC at under 800 MMACs, which frames how much accuracy headroom remains for extreme low-power object detection.","The open-challenges section singles out energy-efficient SNN-based detectors, high-resolution input handling, and transformer-based architectures as the next targets for TinyML object detection.","If the abstract's claimed 90% and 50% transmission reductions were replicated, hierarchical split learning would become a practical bandwidth-saving option for federated training with heterogeneous edge clients."],"supporting_citations":[{"why":"supplies the MCUNetV1-YOLOv2 baseline in Table 8 (51.4% mAP on VOC).","marker":"[88]"},{"why":"supplies the MCUNetV2-YOLOv3 M4/H7 entries in Table 8 (64.6%/68.3% mAP).","marker":"[87]"},{"why":"supplies the fully quantized TinyissimoYOLO entry in Table 8 (56.4% mAP on VOC).","marker":"[110]"},{"why":"supplies the EtinyNet-SSD entry with adaptive scale quantization in Table 8 (56.4% mAP).","marker":"[162]"},{"why":"supplies the XiNet-YOLOv7 S/M/L entries in Table 8 (54.0–74.9% mAP on VOC).","marker":"[4]"},{"why":"supplies the COCO benchmark used for the lightweight OD model comparison in Table 5 and the NAS comparison in Table 6.","marker":"[92]"},{"why":"supplies the PASCAL VOC benchmark used for all MCU-detector comparisons in Table 8.","marker":"[38]"},{"why":"supplies the MLPerf Tiny benchmark context that motivates person detection and OD on MCUs.","marker":"[6]"},{"why":"supplies the DetNAS evolutionary-search baseline in Table 6 (42.0 AP50:95).","marker":"[23]"},{"why":"supplies the EAutoDet gradient-based NAS results in Table 6 (AP50:95 up to 49.2).","marker":"[155]"}],"fun_headline_variants":["TinyML survey disguised as SHeRL-FL paper","Paper titled SHeRL-FL hides TinyML survey inside","Promised split-learning results missing; survey found instead","SHeRL-FL title, but body is a TinyML review","Missing experiments: paper is actually a TinyML survey"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise for the abstract's central claim is that the manuscript actually contains the SHeRL-FL method and its experiments; the full text instead is a different survey, so the claimed 90% and 50% transmission reductions rest on a document that is not present.","fun_headline_variants_meta":{"raw":{"variants":["TinyML survey disguised as SHeRL-FL paper","Paper titled SHeRL-FL hides TinyML survey inside","Promised split-learning results missing; survey found instead","SHeRL-FL title, but body is a TinyML review","Missing experiments: paper is actually a TinyML survey"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000439,"raw_usage":{"total_tokens":2111,"prompt_tokens":832,"completion_tokens":1279,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":576,"completion_tokens_details":{"reasoning_tokens":1195}},"tokens_in":576,"tokens_out":1279,"duration_ms":10415,"temperature":1.0,"reasoning_tokens":1195,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T21:58:04.181403+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader can settle the mismatch immediately by searching the full text for 'SHeRL-FL'; it appears nowhere in the body, and no algorithm, no CIFAR/HAM10000/ISIC experiment, and no transmission-volume table is reported, so there is nothing in the manuscript that could confirm or refute the abstract's numbers.","supporting_citations":[{"cited_title":"DetNAS: Backbone Search for Object Detection","cited_arxiv_id":"1903.10979","evidence_quote":"supplies the DetNAS evolutionary-search baseline in Table 6 (42.0 AP50:95)."},{"cited_title":"EAutoDet: Efficient Architecture Search for Object Detection","cited_arxiv_id":"2203.10747","evidence_quote":"supplies the EAutoDet gradient-based NAS results in Table 6 (AP50:95 up to 49.2)."}],"review_version":1}