{"id":"07029222-b289-4bf9-8094-1745f5de436a","arxiv_id":"2508.12637","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"HOMI is an end-to-end event-camera AI platform achieving 94% DVS Gesture accuracy and 1000 fps throughput on a Xilinx Zynq UltraScale+ FPGA with 33% LUT utilization.","lead":"HOMI is a new edge AI platform that pairs a Prophesee IMX636 event camera with a Xilinx Zynq FPGA to process visual events at up to 1000 frames per second. It reports 94% accuracy on the DVS Gesture dataset while using only about a third of the FPGA's logic resources.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No methodology is present: the 94% accuracy and 1000 fps claims are unverifiable because the submission body is a different paper's full text, leaving no protocol, comparison, or artifact to check.","rationale":"The reader's verdict of UNVERDICTED is correct, and the reader's weakest assumption — that the 94% accuracy depends on an unreported evaluation protocol — is the central issue. My review confirms the full text provided is a different paper's content, which strengthens the reader's concern: there is literally no methodology to inspect. I considered whether the resource claim (33% LUT) could serve as a partial verifiable artifact, but it is not tied to a defined model or configuration, so it does not salvage verifiability. I also note the abstract's internal consistency: the described dual-mode preprocessing (constant-time and constant-event) is plausible and the 1000 fps figure is not inherently implausible for an FPGA accelerator, so I do not claim the numbers are wrong; I only claim they are unsupported. The concrete test I propose is to obtain the real HOMI paper and verify the evaluation protocol and end-to-end measurement definition. If those details are present and standard, the verdict could move toward ACCEPT or CONDITIONAL, but from the current evidence UNVERDICTED is the only honest status.","tokens_in":8069,"tokens_out":1044,"duration_ms":10140,"concrete_test":"Obtain the actual HOMI manuscript (or its supplementary materials) and check whether it reports: (1) a DVS Gesture accuracy evaluation using the standard leave-one-subject-out protocol with a fixed event window (e.g., 1 ms or 5 ms) and class-balanced test set, and (2) an end-to-end throughput measurement that includes sensor readout and preprocessing, not just accelerator inference. If the paper does not contain these details, the headline claims remain unverifiable and the verdict should stay UNVERDICTED.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The central claim — 94% accuracy on DVS Gesture and 1000 fps throughput on the HOMI platform — cannot be evaluated because the manuscript body supplied for review is the full text of an unrelated paper (edgeVLM, arXiv:2508.12638), not the HOMI paper. The HOMI abstract states these numbers but provides no evaluation protocol: no description of the DVS Gesture train/test split, event-window length, class balance, preprocessing parameters (constant-time vs constant-event histogram modes), model architecture, or comparison against published baselines. Without this information, the accuracy figure is not connected to any measurable procedure, so the claim is not falsifiable from the available text. The throughput figure of 1000 fps is similarly ambiguous: it is unclear whether this is sustained end-to-end throughput including event readout and preprocessing, or only accelerator inference throughput for a particular event density. The 33% LUT utilization is a concrete resource figure, but it is not anchored to a defined model size or configuration, so it does not by itself validate the system. Because the central argument rests entirely on headline numbers with no supporting derivation or measurement description, the appropriate status is UNVERDICTED: the claims may be true, but they are not currently verifiable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The abstract of arXiv:2508.12637 describes HOMI, an end-to-end edge AI platform for event cameras combining a Prophesee IMX636 sensor with a Xilinx Zynq UltraScale+ MPSoC FPGA, and reports 94% accuracy on the DVS Gesture dataset, 1000 fps throughput in a low-latency configuration, and 33% LUT utilization. The manuscript body provided for review, however, is the full text of an unrelated paper, 'edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer' (arXiv:2508.12638), which addresses vision-language model inference and contains no mention of event cameras, HOMI, the FPGA implementation, or the DVS Gesture experiments. As submitted, the paper provides no methodology, architecture description, evaluation protocol, or comparisons to support the abstract's claims, making them unverifiable from the available text.","tokens_in":8271,"tokens_out":4189,"duration_ms":39822,"significance":"If the abstract's claims are substantiated, HOMI would represent a useful system-level contribution to low-latency edge processing of event-camera data: a commercial sensor-FPGA pairing with hardware optimized preprocessing, support for both constant-time and constant-event histogram accumulation, and linear/exponential time surfaces would have practical value, and the reported 33% LUT utilization suggests headroom for further integration. However, because the submitted body is a different paper entirely, the significance cannot be assessed beyond the abstract. No code, artifacts, or reproducible protocols are provided, and the headline numbers are not connected to any measurable procedure, so the claims are currently unfalsifiable from the submission.","major_comments":[{"comment":"The body of the submission is not the manuscript described by the abstract. All sections, including the references, present the paper 'edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer' (arXiv:2508.12638), which concerns vision-language model inference and contains no description of the HOMI platform, the Prophesee IMX636 sensor, the Xilinx Zynq FPGA design, the event preprocessing pipelines, or the DVS Gesture experiments. Consequently, the submission provides no technical content by which the abstract's claims for HOMI can be evaluated.","section":"Full text (all body sections)"},{"comment":"The 94% accuracy on the DVS Gesture dataset is asserted without any evaluation protocol. The abstract does not specify the train/test split, event-window length, histogram mode (constant-time versus constant-event), time-surface decay constants, network architecture, or any comparison with published baselines on this dataset. A bare headline accuracy is not a checkable claim.","section":"Abstract"},{"comment":"The throughput figure of 1000 fps is ambiguous because it is not defined whether this is sustained end-to-end throughput including event sensor readout and preprocessing, or inference-only throughput at a particular event density; the batch size and event density are unspecified. Likewise, the 33% LUT utilization is not anchored to a defined model size, bit width, or configuration, so it cannot be interpreted meaningfully.","section":"Abstract"},{"comment":"The abstract claims the platform is 'end-to-end' and 'ultra-low latency,' but no latency measurements, power measurements, or comparisons to prior event-camera gesture-recognition systems or edge AI accelerators are reported. Without baselines, the claimed accuracy and throughput cannot be placed in context, and the statement that the platform 'cater[s] to both accuracy-driven and low-latency applications' is unsupported by any measurement or ablation.","section":"Abstract"}],"minor_comments":[],"recommendation":"reject","confidential_remarks":"The submission package appears to contain the wrong full text: the abstract describes HOMI, while the body is the edgeVLM paper (arXiv:2508.12638). This is a severe verifiability and integrity issue; the editor may wish to confirm the intended submission before any further processing. The manuscript as supplied cannot be reviewed as a research paper because the central claims have no supporting content."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: the HOMI abstract describes a plausible, useful edge-AI platform for event cameras, but the manuscript body is someone else's paper (edgeVLM), so there is no methodology, no baselines, and no way to verify the headline numbers. I can't judge the work on its merits, but what the abstract promises is worth a proper look if the real paper exists.\n\nWhat's actually new: a full sensor-to-FPGA pipeline with hardware-optimized preprocessing (constant-time and constant-event histogram accumulation, linear and exponential time surfaces) and an in-house accelerator, reporting 94% on DVS Gesture and 1000 fps throughput at 33% LUT utilization. That is a reasonable engineering target, and the 94% figure is in the ballpark of published event-camera gesture results, so it's not implausible. The \"end-to-end\" claim and the dual accuracy/latency modes are genuinely useful for robotics people.\n\nThe soft spots are large, but mostly because of the submission problem. The accuracy number has no protocol: no event-window length, no train/test split, no model architecture, no comparison to prior FPGA event-camera systems. The 1000 fps could mean accelerator-only throughput or full pipeline; the abstract doesn't say. There is no artifact, no latency breakdown, no power or energy numbers. And the mismatch between abstract and body makes the submission un-reviewable as-is.\n\nIf I take the abstract at face value, the contribution is solid but incremental: another event-camera FPGA platform, likely with some clever preprocessing, but no evidence of a breakthrough. The absence of any comparison to earlier event-camera accelerators (there is a small literature) is a real missing piece, not just a formatting issue.\n\nMy recommendation: desk-reject this submission, but invite the authors to resubmit with the actual HOMI paper. The abstract alone shows enough engineering substance to warrant a serious referee once the full text matches. If the real paper has the methodology and comparisons, it could be a useful addition to the event-camera systems literature. As it stands, I can't honestly verify anything.\n\nFor your reading group: maybe, as a cautionary tale about arXiv metadata and the importance of self-contained submissions. I wouldn't cite it yet.","headline":"The HOMI abstract promises a credible event-camera FPGA platform with plausible numbers, but the submitted body is an unrelated VLM paper, so nothing can be verified as-is.","tokens_in":8805,"tokens_out":2869,"would_cite":false,"duration_ms":28436,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HOMI is an ultra-low-latency edge AI platform that couples a Prophesee IMX636 event camera with a Xilinx Zynq FPGA, and the paper claims it reaches 94% accuracy on DVS Gesture and 1000 fps throughput while using only 33% of available LUT…","keywords":["event cameras","edge AI","FPGA acceleration","DVS Gesture","ultra-low latency","event-based vision","AI accelerator","hardware-optimized preprocessing"],"falsifier":"Run HOMI on the full DVS Gesture test set with the commonly used event-window length and all ten classes balanced; if the resulting accuracy falls materially below 94%, or if the 1000 fps throughput cannot be sustained on the stated FPGA without losing events, the central claim would be contradicted.","tokens_in":7888,"feed_emoji":"⚡","tokens_out":5305,"duration_ms":46514,"temperature":0.7,"pith_summary":"HOMI is an end-to-end edge AI platform that combines a Prophesee IMX636 event camera with a Xilinx Zynq UltraScale+ FPGA and a custom AI accelerator. The paper reports that, in its high-accuracy configuration, the platform achieves 94% accuracy on the DVS Gesture dataset, and in its low-latency configuration it sustains 1000 frames per second. The aim is to show that event-based perception can be made both accurate and fast enough for closed-loop control entirely on edge hardware, by exploiting the sparsity and asynchrony of event streams through hardware-optimised preprocessing. A sympathetic reader would care because earlier event-processing systems were typically partial, high-latency, or did not fully exploit event sparsity, and HOMI positions itself as a complete alternative.","feed_headline":"Event-camera edge AI hits 94% accuracy at 1000 fps","feed_subtitle":"Pairing a Prophesee sensor with an FPGA, the platform uses only a third of the chip's logic.","key_machinery":"The load-bearing mechanism is the platform itself: a Prophesee IMX636 sensor streaming asynchronous events into a Xilinx Zynq UltraScale+ FPGA that runs an in-house AI accelerator, together with hardware-optimised preprocessing pipelines. These pipelines support two histogram-accumulation modes—constant-time and constant-event—and both linear and exponential time surfaces, converting sparse event streams into regular representations the accelerator can process. The dual-mode design is what lets a single implementation serve both accuracy-driven and low-latency applications, and the claimed 33% LUT usage is the concrete evidence that the design exploits event sparsity rather than over-provisioning compute.","core_discovery":"Stated on the paper's own terms, the central claim is that a complete event-camera pipeline—from sensor to pre-processing to a dedicated neural accelerator on the FPGA—can deliver both competitive recognition accuracy and very high throughput on a single edge device. The authors report 94% accuracy on DVS Gesture in a high-accuracy mode and 1000 fps throughput in a low-latency mode, with the entire hardware-optimised pipeline occupying only 33% of the FPGA's look-up tables. That resource figure is part of the claim: it is offered as evidence that the design leaves substantial headroom for more complex models, multi-task deployments, or further latency reduction, rather than consuming the full device.","pith_inferences":["Since the preprocessing pipelines are presented as general-purpose, HOMI likely extends to other event-based tasks such as optical flow or object tracking, but the paper only demonstrates DVS Gesture as a use case.","If the 94% figure holds under a standard protocol, it would place HOMI at parity with frame-based edge gesture recognizers while gaining the microsecond-level response of event sensors—something frame-based systems cannot match.","A natural next measurement the paper leaves implicit is energy per inference, which would determine whether the platform truly fits battery-powered robots."],"forward_implications":["Gesture-based human-robot interaction and other closed-loop control tasks could run entirely on an event-camera edge device, removing cloud round-trips.","The same platform can switch between accuracy-driven and low-latency modes via its dual preprocessing paths, so one hardware design covers different application needs.","With only 33% of LUTs used, the remaining FPGA resources can host larger models, additional tasks, or further latency optimisation.","The reported figures suggest event-camera edge processing can be competitive with software-based recognizers while running at far higher rates."],"supporting_citations":[],"fun_headline_variants":["Edge event-camera AI: 1000 fps at 94% accuracy","FPGA event-camera AI: 94%, 1000 fps, 33% logic","Event-camera edge AI fits in a third of FPGA","Ultra-fast event-camera AI: 1000 fps, 94% accuracy","Event-camera AI: 94% accuracy, 1000 fps on one FPGA"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline 94% accuracy assumes the DVS Gesture evaluation used a protocol comparable to prior benchmarks—standard event-window length and balanced classes—but the abstract does not report the protocol.","fun_headline_variants_meta":{"raw":{"variants":["Edge event-camera AI: 1000 fps at 94% accuracy","FPGA event-camera AI: 94%, 1000 fps, 33% logic","Event-camera edge AI fits in a third of FPGA","Ultra-fast event-camera AI: 1000 fps, 94% accuracy","Event-camera AI: 94% accuracy, 1000 fps on one FPGA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00069,"raw_usage":{"total_tokens":3120,"prompt_tokens":932,"completion_tokens":2188,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":2083}},"tokens_in":548,"tokens_out":2188,"duration_ms":14597,"temperature":1.0,"reasoning_tokens":2083,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:19:46.871751+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run HOMI on the full DVS Gesture test set with the commonly used event-window length and all ten classes balanced; if the resulting accuracy falls materially below 94%, or if the 1000 fps throughput cannot be sustained on the stated FPGA without losing events, the central claim would be contradicted.","supporting_citations":[],"review_version":2}