{"id":"3f195358-8481-4e80-88b0-7b0f4f2f2c6d","arxiv_id":"2508.01361","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"The paper reports a vision-language-action model fine-tuned for a drone with mid-air haptic actuators, achieving 56.7% target acquisition success and 35-70% generalization across four task axes.","lead":"VLH is a drone that adds a new sense to flying: small mechanical arms that push and vibrate your hands in mid-air, while the drone navigates by seeing the scene and reading your text commands. The project fine-tunes a general-purpose robot AI model to output both flight directions and these touch cues, showing such a combined system can work in real flights.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's '100% accuracy in texture discrimination' is unsupported by the reported experiments, which measure human haptic pattern recognition, not VLH model output.","rationale":"The reader's weakest assumption was the synchronization and alignment of the egocentric VR frame and top-down drone frame (Section 3.2). While that is a real concern about input data quality, it is not the most load-bearing issue: even a perfectly calibrated dual-camera system would not substantiate the paper's central haptic claim, because no experiment reported actually tests whether the VLH model's haptic outputs (Hx, Hy, Hz, Hv) are correct for a given texture, shape, or language command. The label 'texture discrimination' in the abstract and conclusion is not backed by a model-level evaluation; the only discrimination study (Section 5.1) is a human-perception user study. This mismatch is internal to the manuscript and directly undermines the core novelty—that haptics are generated as a consequence of visual and language understanding. The flight results (56.7% target acquisition, 21.3 s, 0.24 m pose error) are plausible and support the navigation aspect, but they say nothing about haptic fidelity. The generalization numbers in Section 5.2 are also ambiguous because 'success' is undefined; if they did include haptic correctness, the paper does not show it. I therefore disagree with the reader's identification of the weakest assumption: the calibration issue would affect both navigation and haptics, but the absence of any model-level haptic metric is a more fundamental gap. Despite this, the verdict remains CONDITIONAL rather than REJECT because the concerns are addressable with additional evaluation, and the reader's verdict already reflects that conditionality. My recommendation is UNCHANGED.","tokens_in":8183,"tokens_out":2475,"duration_ms":31470,"concrete_test":"Re-read Section 5.1 and verify that the '100% accuracy in texture discrimination' does not appear in any table or analysis; the confusion matrices are human responses. Then run a held-out texture classification test: fine-tune VLH as described, hold out one texture category from the 450 scenarios, prompt the model with the texture name, and compare its inferred (Hx, Hy, Hz, Hv) to the ground-truth haptic pattern, reporting per-texture accuracy with a confusion matrix. If the abstract claim cannot be reproduced from the reported experiments or from such a test, remove or rephrase the claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.1 is explicitly a user study: twelve human participants judged nine haptic patterns (three shapes × three vibration levels) rendered by the device. Tables 1–3 and the ANOVA describe human perception, not the VLH model. Yet the abstract and conclusion claim '100% accuracy in texture discrimination' by VLH, and Section 5.3's flight experiments measure only target acquisition (success rate, reach time, pose error), never the correctness of the haptic command (Hx, Hy, Hz, Hv) relative to the requested object/texture. Section 5.2 reports aggregate generalization percentages (70.0/54.4/40.0/35.0) but gives no definition of 'success' and no per-axis breakdown that isolates haptic accuracy. Therefore the central claim that VLH 'co-evolves haptic feedback with perceptual reasoning and intent' rests on a metric that the reported experiments do not measure. Even if the dual-camera frames were perfectly aligned, as the reader's weakest assumption requires, the paper would still lack evidence that the model's haptic outputs are correct. This is a claim–evidence mismatch in the core contribution, not a peripheral calibration issue.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents VLH, a vision-language-action model fine-tuned from OpenVLA via LoRA on a custom dataset of 450 multimodal scenarios. The model takes two camera views (an egocentric VR camera and a top-down drone camera) plus a natural language command and outputs a 7D action vector (Vx, Vy, Vz, Hx, Hy, Hz, Hv), intended to unify drone velocity control with mid-air haptic feedback delivered by a quadcopter equipped with dual inverse five-bar linkage arrays. The reported evaluation consists of (i) a user study in which twelve participants recognized nine haptic patterns (three shapes × three vibration levels), (ii) generalization tests across visual, motion, physical, and semantic axes with reported percentages 70.0%, 54.4%, 40.0%, and 35.0%, and (iii) flight experiments reporting a 56.7% target-acquisition success rate, a mean reach time of 21.3 s, and a mean pose error of 0.24 m. The abstract and conclusion additionally claim 100% accuracy in texture discrimination.","tokens_in":8449,"tokens_out":3712,"duration_ms":44954,"significance":"The idea of treating haptic feedback as a generative output of a vision-language-action model, rather than as a pre-programled reactive channel, is timely and potentially relevant to aerial human-robot interaction and VR. The hardware platform—a quadcopter with two inverse five-bar linkage arrays—is a concrete engineering contribution, and the real-robot flight experiments go beyond pure simulation. However, the significance of the work is currently undermined by a central claim–evidence mismatch: no reported experiment measures the correctness of the model's haptic outputs. The user study in Section 5.1 measures human perception of haptic patterns, not the VLH model's ability to produce correct haptic commands from visual and language inputs. Because the paper's core novelty is the haptic channel, this gap is load-bearing. The generalization and flight metrics also lack precise definitions and statistical support. If the authors add a model-level haptic evaluation and tighten the metric definitions, the contribution could become solid; as it stands, the paper does not support its central claims.","major_comments":[{"comment":"The abstract and conclusion claim '100% accuracy in texture discrimination' by VLH, but Section 5.1 is a human perception study: twelve participants recognized nine haptic patterns, and Tables 1–3 report human confusion matrices and an ANOVA on human responses. No experiment in the paper measures the VLH model's haptic outputs (Hx, Hy, Hz, Hv) against ground truth. This claim must be either supported by a model-level evaluation that compares model outputs to known object/texture labels, or removed.","section":"Abstract, Section 5.1"},{"comment":"The generalization percentages (visual 70.0%, motion 54.4%, physical 40.0%, semantic 35.0%) are reported without any definition of 'success', without per-task or per-trial counts, and without confidence intervals or statistical tests. It is also unclear whether success requires correct haptic output, correct flight action, or both. This is load-bearing because generalization is a core claim of the paper; please define the metric, report the number of trials per axis, and provide confidence intervals or raw data.","section":"Section 5.2"},{"comment":"The flight success rate of 56.7% is defined only as 'the percentage of flights that reached the target area within an acceptable threshold', but the threshold value is never specified. Figure 6 mentions 'stable hovering for at least 5 sec' for green trajectories, yet the text does not state the hover threshold or the success criterion used for the reported 56.7%. The paper also provides no baseline (e.g., OpenVLA without haptic output, a waypoint controller, or random actions) against which this success rate can be interpreted. Please specify the threshold and add a baseline comparison.","section":"Section 5.3"},{"comment":"The model inputs are described as 'two synchronized top-down frames ... both aligned in position and orientation', but the paper reports no calibration procedure, synchronization mechanism, or alignment error between the real-world flight camera and the VR camera. Misalignment between these frames would systematically corrupt the learned mapping from pixels to drone velocities and haptic commands. Please report how the frames were calibrated and synchronized, and quantify the alignment error.","section":"Section 3.2"},{"comment":"The training dataset is said to pair real drone trajectories with VR frames and 'haptic signals (Hx, Hy, Hz, Hv)', but there is no description of how these ground-truth haptic labels were generated or validated. Without a defined reference standard for the haptic signal, it is impossible to assess whether the model's haptic outputs are correct. Please specify the label-generation procedure and, ideally, evaluate haptic outputs against an external benchmark or ground-truth haptic recordings.","section":"Section 4"}],"minor_comments":[{"comment":"The heading contains a typo: 'Traning' should be 'Training'.","section":"Section 4 heading"},{"comment":"The term 'UA V' is inconsistently spaced; it should be written as 'UAV'.","section":"Throughout"},{"comment":"The confusion-matrix header is garbled ('circle square cone%' and 'Answers (Predicted Class)' with misaligned rows); please reformat the table so that actual and predicted classes are clearly labeled.","section":"Table 1"},{"comment":"The ANOVA description states that 'the interaction between vibration and temperature was not significant'; this should read 'vibration and shape'.","section":"Section 5.1"},{"comment":"The contributions list says 'empirical validation through real-world experiments demonstrating 57% positional alignment', which differs from the abstract's '56.7% success rate for target acquisition'. Please reconcile these numbers and clarify what 'positional alignment' means.","section":"Introduction"},{"comment":"The text states that 'Helix represents a generalist VLA model', but the cited reference [16] is about serving large language models over heterogeneous GPUs, not about a VLA model. Please cite the correct Helix robotics model or revise the sentence.","section":"Section 2, Reference [16]"}],"recommendation":"major_revision","confidential_remarks":"The claim–evidence mismatch is serious enough that I would not accept the paper in its current form. The good news is that the gap is addressable: the authors can add a model-level haptic-output evaluation, define metrics precisely, provide calibration details, and either substantiate or retract the '100% texture discrimination' claim. I therefore recommend major revision rather than rejection. The paper may also be better positioned with a stronger related-work comparison to recent VLA-based drone works and with an actual benchmark for haptic generation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe new thing in this paper is real: VLH is the first system I know that takes a vision-language-action backbone (OpenVLA), fine-tunes it on a small bespoke dataset of 450 flight/haptic scenarios, and outputs a joint 7D action vector covering both drone velocity and four haptic channels, all in real time on a flying platform. That integration is genuinely novel relative to VRHapticDrones, HapticGen, RaceVLA, and CognitiveDrone, and the authors did build the hardware and run it. Ninety flights, a 56.7% target-acquisition success rate, and four-axis generalization results are concrete engineering evidence. The paper earns credit for that.\n\nThe soft spots are in the claims, not the build. The abstract says '100% accuracy in texture discrimination,' but Section 5.1 is a user study with twelve humans judging nine rendered haptic patterns. That measures human perception of the device's output, not the VLH model's ability to pick the correct haptic response. Section 5.3's flight tests measure only target acquisition. Nowhere in the paper is Hx, Hy, Hz, Hv compared to ground truth or to a human-judged correctness score for the requested texture. So the central claim that haptics co-evolve with perception and intent is not backed by the reported experiments. The stress-test note is right.\n\nThe generalization numbers (70/54/40/35) are also thin: no definition of success for those tasks, no confidence intervals, no per-axis or per-task counts, and no baseline comparison. The dual-camera input is described as 'aligned in position and orientation' (Section 3.2), but no calibration procedure or alignment error is given, so the reader's weakest assumption is a genuine concern, though secondary to the missing haptic-correctness evidence. All these gaps are addressable: run a study that scores VLH's haptic outputs against ground truth or human ratings, define the hover threshold, add a baseline like OpenVLA without haptics, and report per-task breakdowns.\n\nWho is this for? People working on aerial haptics, VR interfaces, and VLA-controlled drones. It is an application-level integration paper—no new algorithm, no new theory—and its impact is contained to HRI and VR haptics. That is a legitimate niche, just not a transformative one.\n\nMy recommendation: send it to peer review rather than desk-rejecting. The referee process should push for the missing evidence or make the authors soften the claims. With the abstract fixed and the haptic-output evaluation added, this would be a reasonable systems paper. As submitted, it overstates what the experiments show.","headline":"Real novelty in the VLA-plus-aerial-haptics integration, but the abstract's 100% texture accuracy is a user-study number, not a VLH model result—send it to referees, not desk reject.","tokens_in":8995,"tokens_out":2798,"would_cite":false,"duration_ms":31868,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single model turns two camera views and a language command into both drone flight and haptic output.","keywords":["vision-language-action model","haptic feedback","aerial robotics","virtual reality","LoRA fine-tuning","human-robot interaction","drone control","multimodal learning"],"falsifier":"A concrete test is to deliberately offset the two input frames by a known translation and measure whether target-acquisition success stays at 56.7%; if it does, the synchronization premise is not load-bearing. A second check is to rerun the texture-discrimination protocol and see whether the claimed 100% accuracy reproduces.","tokens_in":8001,"feed_emoji":"🛸","tokens_out":9547,"duration_ms":110241,"temperature":0.7,"pith_summary":"This paper is trying to establish that haptic feedback can be generated on the fly by a vision-language-action model rather than being hard-coded as a reactive channel. It fine-tunes a 7-billion-parameter OpenVLA model on 450 scenarios that pair a virtual-reality view, a top-down drone camera, and a language instruction with both flight velocities and haptic outputs. In real flights the system reaches a target object in 56.7% of trials, with a mean reach time of 21.3 seconds and a mean pose error of 0.24 meters. If the claim holds, touch becomes an expressive, context-sensitive channel that can co-evolve with perception and user intent in human-robot interaction.","feed_headline":"Drone model turns language into flight and touch","feed_subtitle":"Fine-tuned on 450 scenarios, it reached targets in 56.7% of 90 flights while also sending force and vibration.","key_machinery":"The load-bearing object is the 7-dimensional action vector that couples flight and touch in one control loop. The named hardware is the inverse five-bar linkage array: two sets of five-bar mechanisms mounted on the drone that push against the user's hands to create localized force and vibration. The model's job is to map two aligned camera frames plus a language command to that vector, so every perceptual decision carries a haptic consequence. The fine-tuned OpenVLA backbone is quantized to INT8 and served at 4–5 Hz, and the dataset is organized as 450 visual-physical-action combinations across three shapes and three texture categories.","core_discovery":"The paper's central claim is that a single vision-language-action model can treat haptic feedback as a generative output rather than a pre-programmed response. The authors fine-tune OpenVLA, a 7-billion-parameter open vision-language-action model, with LoRA on 450 multimodal scenarios; each scenario pairs an egocentric virtual-reality frame, an exocentric top-down drone frame, and a natural language command with a 7D action vector $(V_x, V_y, V_z, H_x, H_y, H_z, H_v)$. The first three components steer the drone, the next three command directional forces on the drone's dual inverse five-bar linkage arrays, and the last sets vibration intensity. In 90 real flights the system is reported to reach target objects in 56.7% of trials, with a mean reach time of 21.3 s and a mean pose error of 0.24 m, and to generalize at 70.0% (visual), 54.4% (motion), 40.0% (physical), and 35.0% (semantic) on novel tasks. The abstract and conclusion additionally report 100% texture discrimination.","pith_inferences":["The same 7D action interface could transfer to other dual-camera teleoperation settings, such as remote inspection or search, where one operator view and one world view are both available.","A natural next step the paper leaves implicit is to invert the mapping: use the haptic command as a diagnostic signal for what the model believes it is touching.","The reported ordering of generalization results suggests that adding physical and semantic training data will yield larger gains than simply expanding visual variety."],"forward_implications":["Haptic patterns can be learned from data instead of hand-programmed, so new virtual objects can carry touch without manual design.","Flight control and feedback share one learned representation, so the model can adjust touch in the same loop that steers the drone.","INT8 quantization at 4–5 Hz is enough for closed-loop aerial interaction, suggesting that deployment on modest GPU hardware is realistic.","The generalization gap is ordered visual > motion > physical > semantic, so enlarging the dataset along physical and semantic axes should improve the weakest behaviors."],"supporting_citations":[{"why":"Supplies the pretrained OpenVLA backbone and the generalization evaluation framework the paper adapts to haptics.","marker":"[15]"},{"why":"The earlier quadcopter-based VR haptic system that VLH extends by adding language-driven generation.","marker":"[17]"},{"why":"A vision-language-action drone navigation model that motivates combining language with flight but omits haptics.","marker":"[7]"},{"why":"Another VLA UAV benchmark demonstrating real-time cognitive task solving without tactile feedback.","marker":"[8]"},{"why":"The hand-based racing-drone control baseline that lacks haptic feedback, against which VLH's added channel is positioned.","marker":"[18]"}],"fun_headline_variants":["Vision-language-haptics model flies drone and delivers touch","Drone model synthesizes touch from language and vision","One model turns words and sight into flight and haptic cues","Vision-language model adds tactile feedback to drone flight","Drone AI learns to generate touch from language and sight"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The two camera views are assumed to be synchronized and aligned in position and orientation, but the paper reports no calibration procedure, synchronization mechanism, or alignment error.","fun_headline_variants_meta":{"raw":{"variants":["Vision-language-haptics model flies drone and delivers touch","Drone model synthesizes touch from language and vision","One model turns words and sight into flight and haptic cues","Vision-language model adds tactile feedback to drone flight","Drone AI learns to generate touch from language and sight"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000715,"raw_usage":{"total_tokens":3278,"prompt_tokens":1069,"completion_tokens":2209,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":685,"completion_tokens_details":{"reasoning_tokens":2130}},"tokens_in":685,"tokens_out":2209,"duration_ms":18471,"temperature":1.0,"reasoning_tokens":2130,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T05:38:29.780799+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test is to deliberately offset the two input frames by a known translation and measure whether target-acquisition success stays at 56.7%; if it does, the synchronization premise is not load-bearing. A second check is to rerun the texture-discrimination protocol and see whether the claimed 100% accuracy reproduces.","supporting_citations":[{"cited_title":"Hoppe, P","cited_arxiv_id":null,"evidence_quote":"The earlier quadcopter-based VR haptic system that VLH extends by adding language-driven generation."},{"cited_title":"Serpiva, A","cited_arxiv_id":null,"evidence_quote":"The hand-based racing-drone control baseline that lacks haptic feedback, against which VLH's added channel is positioned."}],"review_version":1}