{"id":"e78d2660-d851-4226-8d0f-831059240036","arxiv_id":"2607.25049","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A CVaR-based risk-adjusted Ferrari-Canny margin certifies force closure with probability at least β and better ranks adverse-friction grasp success than nominal epsilon.","lead":"Classical grasp scores assume one fixed friction value and miss grasps that fail when surfaces get slippery. This paper scores grasps on the bad tail of a friction distribution and shows the new score better predicts which simulated grasps survive shakes and lifts.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The closure certificate is internally sound, but its practical force depends on an untested, hand-chosen scalar friction prior and a simulator contact model; the experiments demonstrate ranking under the authors’ simulator settings, not calibrated ≥β execution coverage.","rationale":"The reader’s weakest-assumption discussion identifies the same load-bearing issue and appropriately keeps the paper conditional. I found no decisive flaw in the nesting argument behind Corollary 2 or in the CVaR probability bound as stated. The concern is that the probabilistic language can be read as an execution guarantee even though the distribution and contact model are assumptions supplied by the authors. The simulator evidence is useful comparative evidence—especially the consistent ordering over nominal epsilon—but it cannot calibrate coverage because the friction realizations and contact mechanics are selected within the simulation stack. Hardware validation with independently measured, possibly per-contact, friction variation is therefore the decisive next test. Since the reader already assigned CONDITIONAL for precisely these limitations, I would not change the verdict.","tokens_in":26946,"tokens_out":2990,"duration_ms":125519,"concrete_test":"Run a hardware coverage study on LEAP or Allegro across a fixed set of objects and surface treatments. Independently estimate the execution-time friction law—including variation across contacts—from repeated pull/slip measurements, then execute a pre-registered set of ε^(0.9)-certified and noncertified grasps under sampled surface conditions. Record retention and, where feasible, force-closure status. Compare empirical certified-grasp coverage with the promised 0.9, bootstrapping by object and surface; also recompute the metric using the measured heterogeneous-contact friction law. Coverage materially below 0.9, or loss of the reported ranking advantage, would show that the assumed scalar calibrated prior/model rather than the theorem is the limiting factor.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Corollary 2 appears valid under the paper’s model: because the wrench body nests in the common scalar μ, ε(g,v_β)>0 implies closure for every μ≥v_β, and v_β≤VaR_β(μ) gives at least β probability mass above v_β. The vulnerable step is the interpretation of p_μ as execution-time reality. Both priors are stipulated rather than measured, treat friction as one coefficient shared by all contacts, and evaluate closure through linearized quasi-static Coulomb cones. The adverse β=0.9 prior gives v_β≈0.12, while the dynamic studies use fixed simulated coefficients μ∈{0.2,0.3,0.4}, welded contact pads, and mostly gravity-off shake. Thus they test whether ε^(β) ranks outcomes inside the same Coulomb-style simulation family, not whether real certified grasps achieve the advertised ≥0.9 closure probability under an independently estimated contact-friction law. Heterogeneous contact coefficients, compliance, torsional resistance, pad placement, or dynamics could make real coverage far lower even while the theorem remains formally true. This is an external-validity/model-calibration concern, not an internal inconsistency.","agreement_with_reader":"agree"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The manuscript introduces FIRMGrasp, a CVaR-based family of grasp-quality margins for uncertainty in a scalar Coulomb friction coefficient. It defines a risk-adjusted friction v_β=CVaR_β(μ), assembles the grasp wrench space at v_β, and scores a grasp by its signed inscribed-ball radius ε^(β); it also defines Q_β as the lower-tail CVaR of the Ferrari-Canny margin. The paper establishes wrench-space nesting, monotonicity in β, almost-everywhere differentiability, and conditional probabilistic closure certificates. Empirically, it evaluates 1,599 LEAP/Allegro grasps, cross-dataset Shadow Hand grasps, a risk-adjusted synthesis objective, and Drake shake/lift tests. The reported results show substantial nominal-versus-adverse divergence and better ranking of simulated dynamic outcomes than nominal ε and FRoGGeR's min-weight metric.","tokens_in":27290,"tokens_out":8412,"duration_ms":98753,"significance":"If the claims are appropriately scoped, the work provides a useful and computationally attractive robustness axis for grasp evaluation. Once β and a friction prior are chosen, the metric introduces no fitted predictive parameters, requires essentially one LP per stored grasp, and comes with a non-circular conditional closure certificate. The held-out shake and lift studies test outcomes never used to fit the metric, and the cross-hand, cross-dataset, and synthesis experiments make the value of the new axis plausible. Its practical importance nevertheless depends on obtaining realistic friction priors and contact models; the present dynamic discrimination is comparative and moderate rather than a validated physical coverage guarantee.","major_comments":[{"comment":"The closure certificate is mathematically valid only conditionally on the stipulated scalar, contact-shared friction prior and the quasi-static linearized Coulomb model. Neither N(0.7,0.1²) nor 0.5U(0.7,1.0)+0.5U(0.1,0.3) is calibrated from measurements, and the adverse choice gives v_0.9≈0.12. The dynamic studies use fixed simulated μ∈{0.2,0.3,0.4}, welded pads, and the same Coulomb-style simulator family, so they test ranking—not realized ≥β execution coverage under an independently estimated friction law. The abstract's “calibrated friction distribution” and the certificate claims should be qualified accordingly, prior sensitivity or measured calibration reported, and force closure clearly distinguished from dynamic retention.","section":"Abstract; §IX-D; Theorem 1 and Corollary 2; §X-E"},{"comment":"The empirical superiority claim rests on modest AUC differences over highly clustered data: 302 established grasps span only seven objects, with repeated friction/direction outcomes, while the reported bootstrap is grasp-level. The shake gap is 0.629 versus 0.578 (ℓ*) and 0.534 (εnom), and the pick gap over ℓ* is 0.78 versus 0.75 and acknowledged as nonsignificant. Please define the AUC observation unit, report object-cluster paired bootstrap intervals, and preferably include an intention-to-treat analysis counting contact-transfer failures. Wording such as “decisive” and the abstract's success-ordering claim should be tempered until this is done.","section":"§X-E, Figures 12–13 and Tables VIII–IX"}],"minor_comments":[{"comment":"Both objectives leave the median adverse-prior ε^(β) negative (−0.00119 versus −0.00322), so the study demonstrates reduced friction sensitivity rather than predominantly adverse-tail-certified synthesis. Please give confidence intervals for 72/125 versus 86/127 and explain “matched budget” alongside solve times of 19.3 s and 14.9 s.","section":"§X-B, Table V"},{"comment":"The claim of a.e. differentiability of ε(g,μ) alone is not quite enough to interchange differentiation with the tail expectation in Eq. (13). Please state the needed integrability/Lipschitz or finite-sample regularity conditions, or formulate the result through the sample-average subgradient.","section":"§VII, Theorem 3"},{"comment":"εnom is sometimes evaluated with normalized cone edges while ε^(β) uses unit normal forces. This explains why entries such as ε_N^(0.5) can exceed εnom in Table VI, but the convention should be repeated in the table/figure captions or a common scale used for direct comparison.","section":"§IX-F and Table VI"},{"comment":"For DexGraspNet, fingertip forward kinematics followed by projection onto scaled meshes can change contact locations and normals. Given that only 33.2% certify even under the nominal prior, please add a reconstruction sanity check and avoid interpreting cross-dataset fractions as directly comparable across hands and contact conventions.","section":"§X-D, Table VII"},{"comment":"v_β is estimated from 2×10^5 samples even though the chosen Gaussian and mixture-uniform priors admit closed-form tail means. Supplying those formulas would strengthen the “closed-form analytic” description and improve reproducibility.","section":"§IX-E"},{"comment":"Please explain why the per-object median split produces group sizes 109 and 44, and specify pad dimensions/materials, force limits, shake amplitude/duration, and the precise binary adverse-outcome threshold used for AUC.","section":"§IX-G, Table III, §X-E"},{"comment":"The friction ranges are motivated primarily by a commercial reference chart [41]. A measured tribology reference, or explicit presentation of the priors as illustrative stress tests, would be preferable.","section":"§VI and §IX-D"},{"comment":"Algorithm 1 does not use Theorem 3's gradient; the text should consistently present differentiability as enabling future gradient-based synthesis rather than as a capability exercised by the present pipeline.","section":"§VIII"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful bit is simple: take the usual GWS, replace the fixed friction with the lower-tail CVaR of a scalar prior, score the inscribed radius at that discounted friction, and you get a margin whose positivity certifies force closure with probability at least β under nesting. That is not a deep new theory, but it is the right specialization, and they prove the monotonicity, differentiability, and two closure certificates cleanly.\n\nWhat they do well is the empirical package. On 1,599 LEAP/Allegro grasps, 53% of grasps that nominal epsilon certifies lose closure in the adverse tail; correlation with ε_nom collapses from 0.95 to 0.32; the same pattern shows up on DexGraspNet/Shadow; and ε^(β) ranks shake and gravity-on pick better than ε_nom and FRoGGeR min-weight (AUC ~0.63/0.78 vs 0.53/0.67). The lift split at μ=0.2 (70% vs 25%) is the clearest practical number. Synthesis under a risk-adjusted min-weight also trims the sensitive fraction a bit at matched budget. Circularity is low: the metric never sees the dynamic labels.\n\nSoft spots are real but proportionate. Both friction priors are hand-specified, not measured; friction is one shared scalar; cones are linearized hard-finger; execution uses welded pads in Drake, mostly gravity-off shake. So the certificate is internally sound and the ranking is real inside their sim family—the stress-test is right that this is not yet calibrated ≥β coverage on hardware. Absolute AUCs are moderate, per-object power is thin, and the synthesis objective is milder than the adverse eval prior. No shipped code. None of that breaks the central claim.\n\nThis is for people who already care about analytic grasp metrics and want a friction-volatility axis they can drop on candidate pools. Serious referee time is warranted. I would read it, cite the metric and the 53% headline when discussing robustness under contact uncertainty, and bring it to reading group.","headline":"Clean CVaR specialization of Ferrari-Canny that actually separates friction-sensitive grasps; math holds, evidence is sim-only and prior-dependent.","tokens_in":28002,"tokens_out":530,"would_cite":true,"duration_ms":15572,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A risk-adjusted grasp margin scored on the adverse friction tail certifies force closure with probability at least β and flags grasps that classical epsilon rates as safe but fail when friction drops.","keywords":["dexterous grasping","force closure","Ferrari-Canny metric","Conditional Value-at-Risk","friction uncertainty","grasp quality","risk-sensitive robotics"],"falsifier":"Run the same certified-versus-rejected split on physical hardware with measured contact friction drawn from the paper’s adverse mixture: if grasps with ε(β)>0 do not retain the object under lateral pull at low friction at a clearly higher rate than grasps with ε_nom>0 but ε(β)≤0, the predictive claim fails.","tokens_in":27734,"feed_emoji":"🤖","tokens_out":937,"duration_ms":19025,"temperature":0.7,"pith_summary":"Classical grasp scores treat friction as a single known number, so a grasp that looks force-closed at planning time can slip once the real contact is slipperier. This paper replaces that single number with a distribution and scores each grasp on the Conditional Value-at-Risk of the friction tail: it builds the wrench space at the CVaR-discounted friction and takes the inscribed-ball radius of that body as a risk-adjusted margin ε(β). Whenever that margin is positive, the grasp stays force-closed with probability at least β under the assumed prior. On 1,599 LEAP and Allegro grasps, more than half of the grasps the ordinary Ferrari-Canny margin certifies lose closure in the adverse tail; the new margin also ranks shake and pick success better than the nominal score, and certified-robust grasps hold at 70% under a hard lateral pull at μ=0.2 versus 25% for grasps the nominal margin accepts but the risk margin rejects.","feed_headline":"Risk margin catches grasps classical epsilon wrongly certifies","feed_subtitle":"On 1,599 hand grasps, half lose closure in the friction tail; the new score predicts lift success better.","key_machinery":"The risk-adjusted margin ε(β): the inscribed-ball radius of the grasp wrench space assembled at the CVaR-discounted friction v_β = CVaR_β(μ). Positivity of this radius is the closure certificate; the same construction specializes to the classical Ferrari-Canny epsilon under a point-mass prior.","core_discovery":"The authors show that evaluating the Ferrari-Canny force-closure margin at the CVaR mean of the adverse friction tail yields a single scalar ε(β) that is monotone in the confidence level β, differentiable in the grasp parameters, and carries a probabilistic certificate: ε(β)>0 implies force closure with probability at least β. Under calibrated priors this margin separates friction-sensitive grasps that the nominal epsilon rates as high quality, and it orders realized dynamic retention above both the nominal epsilon and a recent min-weight baseline.","pith_inferences":["The same CVaR construction could be applied to uncertain contact normals or object pose, giving a family of risk margins beyond friction alone.","If the friction prior were estimated online from tactile slip, the certificate could be refreshed during grasp execution rather than fixed at synthesis.","Libraries that already expose min-weight or epsilon could add ε(β) as a cheap post-process on stored wrench generators, turning existing grasp pools into friction-robust rankings."],"forward_implications":["Grasp synthesizers can rank or filter candidate contacts by ε(β) without resimulating every friction sample.","A positive risk margin supplies an explicit probability lower bound on force closure under a stated friction prior.","Differentiability of the CVaR margin opens gradient-based synthesis that optimizes the adverse tail rather than a nominal coefficient.","Object geometry that concentrates friction sensitivity (irregular shapes) can be flagged before execution by the drop of ε(β) below zero."],"fun_headline_variants":["CVaR friction margin flags grasps nominal epsilon misrates as safe","Half of epsilon-certified grasps lose closure in adverse friction tail","ε(β) risk margin orders shake and pick success better than Ferrari-Canny","Friction-tail CVaR score certifies grasps that lift at 70% vs 25% rejects","Monotone differentiable ε(β) gives probabilistic force-closure certificate"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The method assumes a calibrated scalar friction distribution really describes execution-time contact, and that quasi-static hard-finger Coulomb friction in simulation is enough for the certificate to predict physical retention.","fun_headline_variants_meta":{"raw":{"variants":["CVaR friction margin flags grasps nominal epsilon misrates as safe","Half of epsilon-certified grasps lose closure in adverse friction tail","ε(β) risk margin orders shake and pick success better than Ferrari-Canny","Friction-tail CVaR score certifies grasps that lift at 70% vs 25% rejects","Monotone differentiable ε(β) gives probabilistic force-closure certificate"]},"model":"grok-4.5","effort":"low","cost_usd":0.003724,"raw_usage":{"total_tokens":1285,"prompt_tokens":942,"num_sources_used":0,"completion_tokens":88,"cost_in_usd_ticks":37244000,"prompt_tokens_details":{"text_tokens":942,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":255,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":942,"tokens_out":88,"duration_ms":4187,"temperature":1.0,"reasoning_tokens":255,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T02:32:50.008592+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same certified-versus-rejected split on physical hardware with measured contact friction drawn from the paper’s adverse mixture: if grasps with ε(β)>0 do not retain the object under lateral pull at low friction at a clearly higher rate than grasps with ε_nom>0 but ε(β)≤0, the predictive claim fails.","supporting_citations":[],"review_version":1}