{"id":"92765181-3b16-4620-bdbb-c7a61d872349","arxiv_id":"2502.07855","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of lightweight vision-language models for edge deployment, marred by citation errors, self-citation, and a lack of selection methodology.","lead":"This paper surveys techniques for compressing and deploying vision-language models on edge devices, including pruning, quantization, knowledge distillation, and federated learning. It catalogs existing methods and applications, but citation errors, self-references, and a missing methodology undermine its reliability as a reference.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Citation integrity failures undermine the survey's central promise: duplicate arXiv IDs and an uncredited first-person passage mean the reference base cannot be trusted as a comprehensive synthesis.","rationale":"I read the paper as a survey, so the applicable standard is trustworthy synthesis rather than novel experimental correctness. The reader's weakest assumption correctly identifies citation accuracy and attribution as load-bearing. The duplicate arXiv IDs are objective, checkable errors, and the first-person BiMediX passage is further evidence of unverified text assembly. A single typo would not be fatal, but the pattern of duplicate IDs across two pairs plus uncredited wording means the evidence base is unreliable. The stronger claim's 'comprehensive' qualifier requires accurate attribution; that condition is not met. The reader already recommended REJECT, and my analysis supports that judgment, so the verdict remains unchanged. This critique targets the argument's reliability, not the authors' intent, and it is independent of any assessment of novelty or mathematical soundness.","tokens_in":36838,"tokens_out":3648,"duration_ms":30847,"concrete_test":"Run an automated citation audit over the full bibliography: resolve every arXiv ID via the arXiv API, compare the returned title/abstract with the citing sentence and the model name in Table II. Specifically, resolve arXiv:2405.09215 and arXiv:2403.19838 and check whether they correspond to ScreenAI/Xmodel-VLM and LightVLP/EM-VLM4AD respectively; and run a text-overlap check of Section IV.A's BiMediX paragraph against arXiv:2402.13253's abstract to quantify uncredited verbatim reuse. If the IDs mismatch or the overlap exceeds a few contiguous sentences, the survey's reliability premise fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central claim is to be a comprehensive and trustworthy synthesis of VLMs for edge networks. That claim rests on the accuracy of its citations and on the text being the authors' own synthesis. Both fail. In the reference list, [53] (Xmodel-VLM) and [59] (ScreenAI) are both assigned arXiv:2405.09215; [52] (LightVLP) and [54] (EM-VLM4AD) are both assigned arXiv:2403.19838. Since one arXiv ID names a single paper, at least two of these entries cannot point to the stated model. Additionally, Section IV.A's Bilingual Medical Mixture LLM paragraph is written in the first person ('We developed a semi-automated English-to-Arabic translation pipeline...') and closely matches the abstract of [136], BiMediX, without quotation or attribution. The reader cannot tell which model a given technique belongs to, and the survey text is not demonstrably the authors' own synthesis. This is not a stylistic quibble: a survey's value is precisely that readers can trace each claim to a verified source; here the trace is broken. The error pattern suggests references were assembled without checking abstracts or IDs, so the claimed 'comprehensive cycle for extending VLMs from the cloud to the edge' rests on an unreliable evidence base.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript surveys vision-language models (VLMs) for edge networks, covering fundamentals, lightweight model designs, efficient fine-tuning methods, edge deployment challenges, applications, and open problems. It positions itself as a comprehensive treatment of the full cycle from cloud-based VLMs to edge deployment, and it provides a taxonomy together with comparison tables of existing surveys and lightweight VLM models.","tokens_in":37042,"tokens_out":3789,"duration_ms":34369,"significance":"If the survey's claims are trustworthy, the paper would be a useful entry point to a rapidly growing area, especially because of its organization around edge deployment (compression, federated learning, security and privacy) and its comparison of existing surveys. The paper is not a technical contribution with proofs or reproducible artifacts, but it does offer a structured taxonomy and broad literature coverage. However, the value of a survey is entirely dependent on the accuracy and provenance of its citations, and it is in this respect that the manuscript currently falls short.","major_comments":[{"comment":"Duplicate arXiv identifiers are assigned to different models: [53] (Xmodel-VLM) and [59] (ScreenAI) both list arXiv:2405.09215, and [52] (LightVLP) and [54] (EM-VLM4AD) both list arXiv:2403.19838. Since an arXiv ID uniquely identifies one paper, at least two of these entries are necessarily wrong, and the reader cannot determine which technique belongs to which model. This directly undermines the survey's promise of accurate representation of prior work and must be corrected for every entry in Table II.","section":"Table II and References [52], [53], [54], [59]"},{"comment":"The paragraph describing BiMediX is written in the first person ('We developed a semi-automated English-to-Arabic translation pipeline...') and closely matches the abstract of reference [136] without quotation or explicit attribution. This is not an original synthesis by the survey authors; it is an apparent verbatim reuse of the primary paper's wording. The passage must be rewritten as a third-person description with appropriate citation, or the manuscript cannot be considered a reliable secondary source.","section":"Section IV.A, Bilingual Medical Mixture LLM paragraph"},{"comment":"The discussion of dynamic inference and model scaling cites reference [168], which is the survey itself, as the sole support for several claims (e.g., the problems of limited computational resources and unpredictable latency, and the mitigation via modular architectures and resource-aware scheduling). Using one's own survey to support one's own recommendations is circular and provides no external evidence. These claims need to be supported by actual prior work, and the self-citation should be removed.","section":"Section V.B, discussion of dynamic inference and model scaling"}],"minor_comments":[{"comment":"The reference points to a 'consensus.app' URL with a 'utm_source=chatgpt' tracking parameter; this is not a proper scholarly citation and should be replaced with the arXiv identifier of the underlying paper (arXiv:2309.16739).","section":"Reference [7]"},{"comment":"There are numerous typographical and formatting errors, including 'comparision' in the opening sentence and inconsistent spacing around model names (e.g., 'RA VEN', 'ALLaV A').","section":"Section II.C"},{"comment":"The figure contains several misspellings: 'Survelliance', 'Heatlhcare', and 'Sem-supervised learning' instead of 'Semi-supervised learning'; these should be corrected.","section":"Figure 1"},{"comment":"The author attribution for EM-VLM4AD is given as 'Z. Wu et al.' in reference [54], but the same paper is correctly attributed to Gopalkrishnan, Greer, and Trivedi in reference [175]; the metadata should be reconciled.","section":"Reference [54]"}],"recommendation":"major_revision","confidential_remarks":"The reader's report recommends rejection on the basis of citation integrity, and I agree that the problems are serious and load-bearing. However, I see them as correctable within the scope of a revision: duplicate arXiv IDs can be fixed by verifying each entry, the BiMediX passage can be rewritten in the authors' own words with proper attribution, and the circular self-citation can be replaced with external sources. That said, the pattern of errors suggests the entire reference list should be audited before resubmission, and the editor may wish to investigate the apparent uncredited reuse in Section IV.A as a research-integrity matter."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nYou asked about arXiv:2502.07855, the survey on VLMs for edge networks. The short version: the paper has a sensible structure but the reference base is too unreliable to trust, and the reader's concerns reproduce exactly when I read it. Two duplicate arXiv IDs, a self-citation used as support for its own recommendations, and an unattributed first-person passage from the BiMediX abstract are all present in the text.\n\nWhat the paper does well: it gives a broad, readable overview of the field—fundamentals of VLM architecture, compression techniques (pruning, quantization, distillation), efficient fine-tuning, federated learning, and edge-specific security/privacy. The comparison table of earlier surveys is a useful device, and a newcomer could get a decent map of the area from the taxonomy. The figures, while occasionally sloppy in labeling, do convey the intended relationships.\n\nThe soft spots are not minor. In Table II, LightVLP [52] and EM-VLM4AD [54] both carry arXiv:2403.19838, and Xmodel-VLM [53] and ScreenAI [59] both carry arXiv:2405.09215. Since one arXiv ID identifies one paper, at least half of those entries are wrong, and the reader cannot trace a model to its source. In Section V.B, the paper cites itself—reference [168]—to support claims about dynamic inference and model scaling problems and their mitigations. That is circular: the survey is being used as evidence for its own claims. In Section IV.A, the BiMediX paragraph is written in the first person and closely tracks the original abstract, without quotation or attribution. That is copied text, and for a survey whose entire value is accurate representation of prior work, it is disqualifying.\n\nThe central argument—that VLMs can be made edge-deployable through compression and adaptation—is not wrong. But the paper's promise is to be a comprehensive, reliable synthesis, and that promise fails. The errors are correctable in principle, but they are pervasive enough that the current version should not go forward.\n\nMy recommendation: reject in current form. If the authors fix the citations, remove the self-citations, and rewrite the copied passage with proper attribution, a revised version could be worth another look. As it stands, I would not bring it to reading group or cite it. The topic is timely, but the execution doesn't meet the bar for peer review.","headline":"A structurally useful survey undone by citation errors, circular self-citation, and a copied passage; reject in current form.","tokens_in":37526,"tokens_out":4013,"would_cite":false,"duration_ms":34127,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims that vision-language models can be efficiently moved from cloud data centers to edge devices through a coordinated cycle of compression, distributed training, and privacy-preserving deployment, and it sets out to map…","keywords":["vision-language models","edge computing","model compression","quantization","knowledge distillation","federated learning","efficient fine-tuning","privacy-preserving inference"],"falsifier":"The 'comprehensive' claim would be falsified by finding an established edge-VLM method or a widely used lightweight VLM that the taxonomy omits; more concretely, a reader can check Table II and find that the identifier 2405.09215 is claimed for both Xmodel-VLM and ScreenAI, which means the survey's citation layer fails a basic consistency test.","tokens_in":36657,"feed_emoji":"📡","tokens_out":5764,"duration_ms":47998,"temperature":0.7,"pith_summary":"This survey sets out to organise the fast-moving field of vision-language models (VLMs) for edge networks into one coherent picture. Its central claim is that deploying VLMs on resource-constrained devices is not a single trick but a full cycle: compressing the model, choosing where each computation runs, training across distributed edge devices, fine-tuning cheaply, and securing the pipeline against privacy and security threats. The authors argue that no existing survey covers this complete cloud-to-edge cycle with attention to security and privacy, and they aim to fill that gap with a structured overview of techniques, models, applications, and open challenges. A careful reader would care because the practical value of multimodal AI in phones, cameras, drones, and medical devices depends on exactly this kind of efficiency playbook.","feed_headline":"A roadmap for squeezing vision-language models onto edge devices","feed_subtitle":"Compression, quantization, and distributed training are mapped as the bridge from cloud AI to phones and drones.","key_machinery":"The central organising object is the 'comprehensive cycle for extending VLMs from the cloud to the edge', presented as a six-step design process in Figure 6. Each step is a distinct category of methods: (1) data selection and pre-processing, (2) model choice across edge and cloud, (3) distributed implementation using federated learning with parent-child model partitioning, (4) post-processing and evaluation, (5) deployment and continuous learning, plus (0) the compression layer of pruning, quantization, and knowledge distillation that runs through the whole cycle. The survey also uses a three-part architecture distinction from the VLM literature (single-stream vs dual-stream encoders, plus fusion mechanisms) and a task taxonomy in Figure 8 (image-text, video-text, and vision-as-VL tasks). These frameworks carry the survey's argument because they turn scattered techniques into a checklist.","core_discovery":"On the paper's own terms, the discovery is taxonomic: it identifies the load-bearing methods that make edge VLMs viable and arranges them into a deployable pipeline. The survey distinguishes general lightweight VLMs from edge-specific VLMs, presents the six-step design process of data selection, model choice, distributed implementation, post-processing and evaluation, and continuous learning, and catalogues eight open problem areas including compressed lightweight VLMs, context-aware models, cross-modality adaptation, security, privacy, and communication-efficient architectures. It also compiles a catalogue of concrete models (MobileVLM V2, EfficientVLM, MiniVLM, EdgeVL, Moondream2, and others) and applications in healthcare, environmental monitoring, autonomous driving, and surveillance. The paper's contribution is the claim that these pieces fit together into a single comprehensive cycle that researchers and engineers can follow.","pith_inferences":["Because the survey does not reproduce evaluation setups for the models it cites, its specific numbers (for instance accuracy gains or model sizes) should be verified against the primary sources before being used in design decisions.","The survey's own Table II assigns the same arXiv identifier (2405.09215) to two different models, Xmodel-VLM and ScreenAI; that suggests the citation layer is not fully reliable and the map should be re-checked against primary literature.","A natural extension would be a quantitative comparison benchmark that measures latency, memory, and energy of the surveyed compression methods on identical edge hardware; the survey itself provides no such head-to-head numbers.","The privacy discussion points toward a testable design: a federated prompt-learning plus homomorphic-encryption combination for edge VLMs could be evaluated for accuracy loss and communication cost; the survey lists the ingredients but does not run that experiment."],"forward_implications":["Practitioners get a structured checklist for taking a VLM to the edge: compress first, then decide edge vs cloud placement, then distribute training, then evaluate and deploy with continuous learning.","Security and privacy are treated as first-class steps in the deployment cycle, not afterthoughts: differential privacy, homomorphic encryption, secure aggregation, and federated prompt learning are mapped to specific stages.","The survey's taxonomy implies that general lightweight VLMs are not automatically edge-ready; edge-specific constraints such as energy budgets and on-device inference matter.","The eight open challenges named by the survey, such as context-aware VLMs and communication-efficient distributed inference, mark the concrete gaps where future contributions are needed."],"supporting_citations":[{"why":"Foundational pruning/compression technique that the survey's compression layer is built on.","marker":"[9]"},{"why":"Supplies the quantization method described in the compression schemes.","marker":"[10]"},{"why":"Supplies the knowledge-distillation method central to the compression discussion.","marker":"[11]"},{"why":"Supplies the single-stream/dual-stream architecture taxonomy used in Section II.","marker":"[35]"},{"why":"The efficient fine-tuning survey that the paper positions itself against for prompt and adapter methods.","marker":"[36]"},{"why":"The efficient multimodal-LLM survey that the paper compares against in its scope table.","marker":"[41]"},{"why":"A representative lightweight edge VLM whose architecture (LDPv2 projector, MobileLLaMA) is described in detail.","marker":"[50]"},{"why":"The EdgeVL framework used as the main example of edge-specific, cross-modality VLM adaptation.","marker":"[87]"},{"why":"Source of the vision-language task taxonomy (image-text, video-text, vision-as-VL) shown in Figure 8.","marker":"[135]"}],"fun_headline_variants":["Survey charts the route for edge-deployed vision-language models","VLM edge playbook: compress, quantize, deploy","From cloud to edge: how to slim down VLMs","Comprehensive survey on edge-optimized vision-language models","A guide to making vision-language models edge-ready"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole survey rests on the assumption that the roughly two hundred works it cites are accurately described, correctly attributed, and representative of the field; the duplicate arXiv identifier in its own Table II shows that this assumption is not safe.","fun_headline_variants_meta":{"raw":{"variants":["Survey charts the route for edge-deployed vision-language models","VLM edge playbook: compress, quantize, deploy","From cloud to edge: how to slim down VLMs","Comprehensive survey on edge-optimized vision-language models","A guide to making vision-language models edge-ready"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000618,"raw_usage":{"total_tokens":2834,"prompt_tokens":876,"completion_tokens":1958,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":492,"completion_tokens_details":{"reasoning_tokens":1878}},"tokens_in":492,"tokens_out":1958,"duration_ms":13041,"temperature":1.0,"reasoning_tokens":1878,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T12:16:26.021187+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The 'comprehensive' claim would be falsified by finding an established edge-VLM method or a widely used lightweight VLM that the taxonomy omits; more concretely, a reader can check Table II and find that the identifier 2405.09215 is claimed for both Xmodel-VLM and ScreenAI, which means the survey's citation layer fails a basic consistency test.","supporting_citations":[{"cited_title":"Vision-Language Intelligence: Tasks, Representation Learning, and Large Models","cited_arxiv_id":"2203.01922","evidence_quote":"Supplies the single-stream/dual-stream architecture taxonomy used in Section II."},{"cited_title":"Self-Adapting Large Visual-Language Models to Edge Devices across Visual Modalities","cited_arxiv_id":"2403.04908","evidence_quote":"The EdgeVL framework used as the main example of edge-specific, cross-modality VLM adaptation."},{"cited_title":"Vision-Language Pre-training: Basics, Recent Advances, and Future Trends","cited_arxiv_id":"2210.09263","evidence_quote":"Source of the vision-language task taxonomy (image-text, video-text, vision-as-VL) shown in Figure 8."}],"review_version":1}