{"id":"c6fae92e-c8e8-41c6-b60f-35ea9d57df05","arxiv_id":"2411.18979","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A Fin Ray gripper with an optimized multi-mirror optical layout and learned perception achieves large-deformation, single-camera visual-tactile sensing with reported force and proprioceptive accuracy.","lead":"A soft robotic gripper finger uses a camera and an array of mirrors inside a flexible Fin Ray structure, so a single camera can keep seeing the contact surface even while the finger bends deeply. It reports force accuracy of 0.14 N, contact-position accuracy of 0.19 mm, and multimodal sensing of texture, temperature, and slip.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The full-coverage claim rests on an unvalidated 2D mirror-layout optimization; no measured coverage under real three-dimensional or torsional loads is reported, so the central advantage is not established.","rationale":"The reader's weakest assumption identifies the planar 2D force-deformation mapping and the rigid mirror-follows-node assumption as the principal unverified link in the argument. My reading agrees: the full-coverage claim, which is the paper's core novelty, depends on the optimized mirror layout staying valid under real deformation. The paper does not provide a quantitative coverage metric, an independent validation of the optimized layout on the physical sensor, or the optimized layout parameters themselves. This is therefore a load-bearing concern, not a minor reproducibility issue. Other weaknesses exist, such as the single validation split for force accuracy and the uncontrolled comparison of deformation with GelSight Baby Fin Ray, but they are secondary: even if those numbers are accepted, the full-coverage claim would remain unsupported without a direct coverage measurement under realistic loading. The concern does not invalidate the presented prototype or the qualitative demonstrations; it means the central claim should remain conditional on further validation. Since the reader already assigned a CONDITIONAL verdict, no change to the verdict is needed.","tokens_in":15699,"tokens_out":9028,"duration_ms":97770,"concrete_test":"Measure the coverage ratio from segmentation masks of the sensing region under a grid of loads: normal forces from 0 to 7.5 N at multiple x-y locations, oblique loads at ±20°, and torsional loads applied to the fingertip. For each condition, compute the fraction of annotated sensing-region skeleton pixels that are visible in either the direct or mirror-reflected image. If the coverage ratio falls below 95% under any load within the claimed deformation range, or if a continuous blind region appears at the fingertip, the full-coverage claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing weakness is that the central 'full-coverage under large deformation' claim is never validated against the actual three-dimensional deformation of the integrated sensor. Section III-B optimizes the mirror layout from a planar force-deformation map f: F → {N_i, P_i} sampled from a bare Fin Ray with identical structural parameters, parameterizing each mirror angle θ_i relative to the chord between adjacent back-beam nodes and assuming the rigid mirror follows that chord without distortion or separation. The real sensor, however, carries bonded T-shaped mirrors, a camera, LED strips, and a PDMS/silicone pad; these alter the stiffness and local curvature of the back beam, and the mirror orientation is set by local tangent and curvature rather than the nodal chord. Torsional or out-of-plane loads, which are common when grasping curved objects, are outside the two-dimensional model entirely. Section IV provides only qualitative images (e.g., Fig. 6) and no measured coverage ratio as a function of load or deformation. The paper itself admits that the front beam occludes the fingertip region and that mirrors are needed to recover the occluded texture (Sec. IV-B2), yet no coverage fraction is quantified. If the two-dimensional optimization mispredicts mirror poses under real loading, the optimized layout may leave blind regions precisely in the large-deformation grasps the sensor is designed to handle.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Gelsight FlexiRay, a Fin Ray-based soft gripper finger that integrates a single camera, a multi-mirror optical system, a flexible silicone/PDMS tactile pad, and thermochromic markers to provide force, contact position, proprioception, texture, temperature, and slip sensing. The key design idea is to treat large structural deformation as a design input rather than a failure mode: a CMA-ES optimization over mirror angles, positions, lengths, and camera pose is used to maintain optical coverage of the tactile surface under deformation, using a planar force-deformation map of the Fin Ray structure. The authors report a force RMSE of 0.135 N, a contact-position mean error of 0.83 mm, average side-beam node positioning errors around 0.19 mm, a texture classification accuracy of 88.33%, and a cup-transfer human-robot interaction demonstration, and they claim roughly fivefold larger deformation under load compared with existing compliant visual-tactile sensors.","tokens_in":15784,"tokens_out":4197,"duration_ms":45810,"significance":"If the central claims are substantiated, this is a useful step toward high-resolution, large-coverage tactile sensing in compliant grippers: a single camera plus passive mirrors is considerably simpler and cheaper than multi-camera segmented coverage, and the explicit optimization of the optical layout under deformation is a sensible design methodology. The force and proprioception experiments use external ground truth (load cell and global camera), so those evaluations are not circular. The paper also contains real hardware demonstrations, learning-based perception models with reasonable reported performance, and a clear multimodal task set. However, the headline claims of 'full coverage' under large deformation and 'fivefold greater deformation' are not yet supported by quantitative measurements, and the reported accuracy numbers lack uncertainty estimates and repeated-trial validation. The central ideas are publishable, but the evidence base currently falls short of the claims.","major_comments":[{"comment":"The central 'full-coverage under large deformation' claim is not validated on the integrated sensor. The objective function in Eq. (2) maximizes coverage over planar 2D nodes {P_i} obtained from f: F -> {N_i, P_i}, parameterizing each mirror angle relative to the chord between adjacent back-beam nodes and assuming the rigid mirror follows that chord without distortion or detachment. The real sensor carries bonded T-shaped mirrors, a camera, LED strips, and a PDMS/silicone pad, which alter the local stiffness and curvature of the back beam, and torsional or out-of-plane loads are outside the 2D model. Section IV provides only qualitative images (e.g., Fig. 6) and does not report a measured coverage fraction as a function of load or deformation. Please add direct measurements of what fraction of the tactile sensing region remains visible under controlled normal, shear, and torsion-like loads on the integrated finger, and state the resulting optimized layout parameters so the optimization can be independently assessed.","section":"III-B, Eq. (2)"},{"comment":"The quantitative accuracy claims are based on a single training/validation split without repeated trials or confidence intervals. The force RMSE of 0.135 N, correlation of 0.997, contact-position mean error of 0.83 mm, and average node positioning error of about 0.19 mm are reported from one split of 5,000 images at a 4:1 ratio. Deep-network results can vary with random initialization and data split, so these point estimates do not establish the claimed accuracy levels. Please report cross-validated results, repeated-seed statistics, or confidence intervals for the force and positioning metrics, and clarify the exact metric being called 'positioning accuracy' in the abstract and conclusion.","section":"IV-A2"},{"comment":"The 'fivefold larger structural deformation under the same loads' comparison is not measured directly. The text states that GelSight Baby Fin Ray shows 'around 3-4 mm' deformation at 7.5 N based on a cited prior sensor, while FlexiRay reaches about 15 mm in the authors' setup, and the conclusion converts this into 'fivefold greater deformation capacity.' Because the contact geometry, probe type, loading protocol, and deformation metric (contact depth versus structural node displacement) are not matched to the baseline experiment, this headline comparison is not established. Please perform a side-by-side measurement with the same loads, probes, and deformation definition, or explicitly label the comparison as qualitative and remove the quantitative factor claim.","section":"IV-A2"},{"comment":"The optimized design parameters that are central to the contribution are not disclosed. Decision variables in Eq. (1) include mirror angles, offset distances, mirror lengths, camera distance coefficient u, and optical-axis angle phi, but the resulting optimized values are not reported anywhere in the manuscript. Without these parameters, the CMA-ES optimization cannot be reproduced or assessed, and the claim that the layout is 'systematically optimized' is not checkable. Please include the optimized parameter set, or provide them in a supplement.","section":"III-B"}],"minor_comments":[{"comment":"The abstract and conclusion state a force accuracy of 0.14 N, while Section IV-A2 reports an RMSE of 0.135 N; please make the metric and rounding consistent, and prefer 'RMSE' over 'accuracy' or define what 'accuracy' means.","section":"Abstract and IV-A2"},{"comment":"The term 'full coverage' is used throughout but never formally defined. A precise definition, such as the fraction of tactile-surface points visible by direct or reflected rays under a given load, would strengthen the paper and make the coverage claim testable.","section":"II and IV"},{"comment":"The ray-coverage radius is introduced as 'R ∝ lc' without a definition of the proportionality constant or of lc; please specify how the indicator function I(x, r, p, R) is evaluated and how many rays m and target points per deformation are used.","section":"III-B, Eq. (2)"},{"comment":"Figure 5(F) shows ten repeated continuous interactions, but the plotted data are shown without error bands or statistical summaries; adding confidence intervals would substantially strengthen the dynamic-force and contact-depth results.","section":"IV-A2, Fig. 5(F)"},{"comment":"The texture-classification accuracy of 88.33% is based on 120 test grasps and includes a large per-class spread (e.g., 73.3% for the cyan ball); reporting confidence intervals and class-wise sample sizes would help the reader judge robustness.","section":"IV-C"},{"comment":"The temperature-sensing experiment demonstrates only discrete discrimination among three cup temperatures and does not report a temperature error, response time, or the number of repeated trials; please state quantitative performance or explicitly describe the experiment as a qualitative demonstration.","section":"IV-D"},{"comment":"There are several typographical and wording issues, including 'neutral network' in the index terms, 'complaint finger framework' in Section III-A, and 'transfering' in Section IV-D; a careful language pass is recommended.","section":"Index Terms and throughout"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the unvalidated 2D mirror-layout model is the main substantive issue: it directly targets the paper's central 'full coverage under large deformation' claim, and it should be addressed with quantitative coverage measurements on the integrated sensor. The other load-bearing issue is the absence of uncertainty quantification for the headline accuracy numbers. These are fixable within the scope of the manuscript, and I do not see grounds for rejection. I would also recommend the authors make the optimized layout parameters available, as the optimization is a core claimed contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, the paper has a real idea: put a multi-mirror reflector system on the flexible back beam of a Fin Ray gripper and optimize the mirror/camera geometry with CMA-ES so a single camera keeps seeing the contact surface while the finger deforms. Prior GelSight Fin Ray designs kept the back beams rigid to preserve the optical path. FlexiRay deliberately mounts the optics on the deforming structure. That is a genuine departure and it is the reason to read the paper. The temperature-sensitive pigment layer and the separation of direct vs mirror-reflected perception into separate segmentation channels are also reasonable engineering choices.\n\nSecond, the quantitative claims are softer than the abstract suggests. The headline 0.14 N force RMSE and 0.19 mm proprioception accuracy come from a single training/validation split, with no confidence intervals or repeated runs. The '5 times larger deformation' line is inconsistent in the text (Section IV-A2 says 'more than four times') and the comparison uses a cited GelSight Baby Fin Ray value (3–4 mm at 7.5 N) rather than a same-setup baseline measurement. The 'full-coverage' claim is softened in the experiment section, where the authors admit the front beam occludes the fingertip region and mirrors only recover part of the contact area; no coverage fraction is ever quantified. The stress-test point you passed along is fair: the mirror layout is optimized on a 2D force-deformation map of the bare Fin Ray, with mirrors assumed to rigidly follow the back-beam chord. Real loads include torsion and out-of-plane bending, and the bonded mirrors/camera/LEDs change the local stiffness. The paper gives qualitative images showing the mirror path works in some grasps, but no measured coverage ratio as a function of load. No code, data, trained models, or optimized layout parameters are released, so exact replication is not possible.\n\nWhat is solid: the force and position evaluations are supervised regressions against external ground truth (load cell, global camera), so the circularity burden is low. The texture classification on eight balls, with a confusion matrix and 88% average accuracy on held-out grasps, is a legitimate demonstration, though the misclassification story is a bit hand-wavy. The temperature and slip experiments are qualitative success demonstrations; no trial counts or error rates are given.\n\nWho is this for: anyone working on vision-based tactile sensors in soft grippers. It is a useful systems paper with a genuinely new optical design. I would not base quantitative claims on it yet—the numbers need a proper uncertainty treatment and a measured baseline comparison. But the design concept is worth citing, and as a prototype demonstration it deserves referee time. My recommendation: send it to peer review with the expectation of heavy revision—measure the deformation baseline in the same setup, report confidence intervals or repeated runs, quantify coverage under 3D/torsional loads, and release the optimized layout parameters.","headline":"A genuinely new multi-mirror-on-flexible-structure VTS design, but the headline numbers and 'full-coverage' claim outrun the current evidence.","tokens_in":16505,"tokens_out":3092,"would_cite":true,"duration_ms":29566,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GelSight FlexiRay claims that one camera, aided by passively reorienting mirrors, keeps full tactile coverage of a Fin Ray soft gripper while it deforms roughly five times more than prior compliant visual-tactile sensors, reaching 0.14 N…","keywords":["vision-based tactile sensor","soft gripper","Fin Ray effect","multi-mirror optics","CMA-ES layout optimization","multimodal tactile sensing","proprioception","compliant gripping"],"falsifier":"Run the same UR5e loading setup but push the hemispherical probe at roughly 30 degrees off the z-axis while recording the internal image; if a noticeable fraction of the sensing-pad markers in the mirrored regions disappears from view or the force-estimation RMSE rises above the normal-force training range, the planar deformation assumption is violated.","tokens_in":15307,"feed_emoji":"📷","tokens_out":6302,"duration_ms":52514,"temperature":0.7,"pith_summary":"The paper sets out to show that a vision-based tactile sensor need not be rigid: mounted on a Fin Ray soft gripper, a single camera and an array of small mirrors can keep the whole contact surface in view even while the finger bends deeply. The authors model the gripper's force-to-deformation behavior, then use CMA-ES to place camera and mirrors so that reflected rays cover the sensing pad across many deformation states. They report force estimation error of 0.14 N and proprioceptive joint-position error of about 0.19 mm, with roughly five times larger deformation under the same load than a prior compliant GelSight Fin Ray design. If true, this would let one inexpensive camera deliver force, contact location, texture, temperature, slip, and finger posture simultaneously on a soft compliant gripper.","feed_headline":"Fivefold deeper bends, one camera still sees the touch","feed_subtitle":"An optimized mirror layout keeps full contact coverage while a Fin Ray finger bends, with 0.14 N force error and 0.19 mm pose error.","key_machinery":"The load-bearing mechanism is a multi-mirror optical relay whose layout is optimized offline from a measured planar force-deformation map $f: F \\to \\{N_i, P_i\\}$ of the Fin Ray. Each T-shaped planar mirror is bonded to the flexible back beam at a small footprint, so as the beam bends the mirror passively reorients with it; the optimizer chooses mirror angles, midpoint offsets, lengths, camera position coefficient $u$, and optical-axis angle $\\phi$ to maximize the number of target points on the sensing pad hit by camera rays, either directly or after one reflection, sampled over $K$ load states, with occlusion and safety penalties. The argument is that this turns structural deformation into a self-adjusting reflector geometry that keeps the contact surface visible to one camera.","core_discovery":"The paper's central claim is that optical occlusion during large structural deformation can be converted from a fatal flaw into a design variable. The authors sample the Fin Ray's node displacements under different loads, feed that force-deformation map into a CMA-ES optimization of mirror angles, offsets, lengths, and camera pose, and obtain a layout under which direct and mirror-reflected views together cover the tactile pad across the deformation range. The resulting TPU-and-PDMS finger with a 12-megapixel wide-angle camera simultaneously estimates contact force (RMSE 0.135 N, reported as 0.14 N accuracy), 3D contact position (mean error 0.83 mm), and 14 proprioceptive joint-node positions (average error about 0.19 mm), classifies textures (95.83% validation, 88.33% average success in random grasps), distinguishes water temperature, and detects slip during handover. It also reports about 15 mm contact-depth deformation at 7.5 N, versus roughly 3-4 mm for the GelSight Baby Fin Ray, which it summarizes as fivefold larger deformation under the same loads.","pith_inferences":["The same CMA-ES mirror-layout recipe should transfer to any soft structure with a measurable or simulable deformation field, not only Fin Ray fingers; a natural next test is a bending or twisting continuum arm.","Because the optimization uses a planar 2D cross-section map, out-of-plane shear or torsion during real grasps is the most likely failure mode; a stress test with off-axis loading would reveal whether the full-coverage claim extends to 3D deformation.","The reported joint-position error grows at the two lower back-beam nodes, where deformation is largest; adding a marker or a small extra mirror aimed at that beam segment could close the gap.","A transparent or highly specular grasped object could confuse the direct-versus-reflected image segmentation, since the mirrored view duplicates the pad; testing such objects would probe the limits of the learning-based decoupling."],"forward_implications":["A single low-cost camera can serve as the entire tactile front end of a soft gripper, with mirrors replacing a second camera for segmented coverage of bent regions.","Because the mirrors move passively with the back beam, the optical system does not stiffen the finger; the reported deformation at 7.5 N is about 15 mm, several times the 3-4 mm of the GelSight Baby Fin Ray.","Force, contact location, joint posture, temperature, texture, and slip are all read from the same image stream, so a gripper can combine compliance with rich feedback without additional skin electronics.","The learned perception models generalize across four probe geometries and dynamic continuous contact, supporting force and depth tracking during active pressing.","The demonstrated sorting and cup-handover tasks indicate the multimodal outputs are usable for object classification and safe human-robot release."],"supporting_citations":[{"why":"Provides the compliant GelSight Baby Fin Ray baseline whose deformation under load (3-4 mm at 7.5 N) is compared with FlexiRay's roughly 15 mm.","marker":"[35]"},{"why":"Prior GelSight Fin Ray integration that kept back beams rigid to preserve the optical path; FlexiRay's design is positioned against this limitation.","marker":"[14]"},{"why":"Supplies the CMA-ES algorithm whose update rules the paper uses to optimize the mirror and camera layout.","marker":"[36]"},{"why":"The Python CMA-ES implementation used to run the layout optimization.","marker":"[38]"},{"why":"PP-LiteSeg semantic segmentation model that extracts the front beam, sensing region, and contact region from raw images.","marker":"[39]"},{"why":"ResNet-style backbone used for the proprioception and texture classification heads.","marker":"[40]"},{"why":"Foundational retrographic/VTS imaging principle that the FlexiRay sensing pad and marker tracking build on.","marker":"[24]"},{"why":"GelSight high-resolution tactile sensor reference that motivates the visual-tactile approach and its force-estimation baseline.","marker":"[10]"}],"fun_headline_variants":["Mirror tricks turn finger bends into full-coverage touch","Optimized mirrors let soft fingers bend deep without losing sight","Fivefold bend, full coverage: GelSight FlexiRay keeps eyes on touch","Deformation no longer blinds soft touch sensors","Fin Ray finger bends five times more, sensor still sees all"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The mirror layout is optimized from a planar 2D force-deformation map, assuming each mirror rigidly follows the back-beam nodes without distortion or detachment; if real grasps introduce torsion or out-of-plane bending, the claimed full coverage may not hold.","fun_headline_variants_meta":{"raw":{"variants":["Mirror tricks turn finger bends into full-coverage touch","Optimized mirrors let soft fingers bend deep without losing sight","Fivefold bend, full coverage: GelSight FlexiRay keeps eyes on touch","Deformation no longer blinds soft touch sensors","Fin Ray finger bends five times more, sensor still sees all"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000224,"raw_usage":{"total_tokens":1535,"prompt_tokens":1096,"completion_tokens":439,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":712,"completion_tokens_details":{"reasoning_tokens":354}},"tokens_in":712,"tokens_out":439,"duration_ms":4467,"temperature":1.0,"reasoning_tokens":354,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:40:07.411381+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same UR5e loading setup but push the hemispherical probe at roughly 30 degrees off the z-axis while recording the internal image; if a noticeable fraction of the sensing-pad markers in the mirrored regions disappears from view or the force-estimation RMSE rises above the normal-force training range, the planar deformation assumption is violated.","supporting_citations":[{"cited_title":"Gelsight baby fin ray: A compact, compliant, flexible finger with high-resolution tactile sensing,","cited_arxiv_id":null,"evidence_quote":"Provides the compliant GelSight Baby Fin Ray baseline whose deformation under load (3-4 mm at 7.5 N) is compared with FlexiRay's roughly 15 mm."},{"cited_title":"Gelsight fin ray: Incorporating tactile sensing into a soft compliant robotic gripper,","cited_arxiv_id":null,"evidence_quote":"Prior GelSight Fin Ray integration that kept back beams rigid to preserve the optical path; FlexiRay's design is positioned against this limitation."}],"review_version":1}