{"id":"901300c2-6e14-4d03-8429-5d0fdfe8441a","arxiv_id":"2411.13148","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"By feeding a policy the remaining time to a speed-based deadline and clipping per-step rotation reward, the authors train in-hand reorientation policies that match target speeds from 0.25 to 1.5 rad/s and reach up to 2.0 rad/s.","lead":"The authors trained AI controllers that let a four-fingered robot hand rotate a cube between its fingers using only touch and joint sensors, with the rotation speed set by a command signal. The system is claimed to be the fastest vision-free in-hand reorientation on a real robot, and it can slow down or speed up on demand.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Real-hardware speed matching rests on subjective video timing of 30 trials; objective pose tracking is needed to support the headline speed-adjustable and 'fastest ever' claims.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern I find: the real-hardware validation of speed-adjustability and the 'fastest ever' comparison depends on subjective video inspection rather than objective measurement. This is the most consequential gap because the paper's headline contributions include zero-shot transfer with speed control to a real robot. The simulation results—speed correlation up to 1.5 rad/s, saturation near 2.0 rad/s, 93.5% success—are credible and well described. However, the real-world experiments are the only evidence that the method transfers and that the speed-control behavior is preserved outside simulation; the paper itself admits the trials were too few for quantitative comparison. I considered alternative concerns, such as the feedforward policy's ability to interpret the remaining-time signal ξ = Td − t without explicit time, but the policy also receives the current relative rotation, so it can infer target speed from ξ together with the current angle; this is not a clear flaw. The time-optimal framing is also not proven optimal, but the paper only claims to push speed with a time-optimal-style objective, which is acceptable. The lack of code and data is a reproducibility limitation, but not the central scientific weak point. Therefore the conditional verdict is appropriate: if objective tracking on the real hand confirms the timing, the central claim is substantially supported; if not, the real-world speed-adjustable claim remains unverified.","tokens_in":8951,"tokens_out":9056,"duration_ms":93143,"concrete_test":"Run the real-hand protocol with an external motion-capture or high-frequency vision system that records the cube pose with sub-100 ms temporal resolution for at least 30 trials per target speed (0.5, 1.0, 2.0 rad/s). Compute the time T from start until d(Rg, R_t) < 0.4 rad for each trial, and compare the measured average speed θ0/T against the commanded ωd. Report per-speed mean, standard deviation, success rates, and the mean absolute relative error |T - Td|/Td. If the mean relative error exceeds about 10% or the speed-matching relationship flattens or becomes non-monotonic at any speed, the real-world speed-adjustable claim is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of zero-shot transfer with speed control to the real DLR-Hand II is supported only by human inspection of video footage. Section V-B states that 'the success of the episode and the time required to reach the goal are determined by examining the video footage' and that 'we could not run enough trials to compare quantitatively.' Only 30 trials were run, and for only 'a few trials' was the timing checked carefully, with '<10% off' claimed for those. At ωd = 2.0 rad/s and a π rotation, the target time is about 1.6 s, so a 10% error is 160 ms; such errors are plausible when judging from video when a cube crosses a 0.4 rad threshold. The 'fastest ever' claim is also comparative and would require a consistent, objective time-to-goal measurement. Without an external pose tracker, the real-world speed-adjustability and the speed-matching accuracy are essentially unquantified for all but a few trials. The simulation evidence (Fig. 5) is credible and supports the method in principle, but it does not by itself establish the real-world transfer of precise speed control; the hardware demonstration is the load-bearing evidence for that part of the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes reinforcement learning objectives for time-optimal and speed-adjustable in-hand reorientation of cuboids with the DLR-Hand II, using only tactile/proprioceptive feedback. The main technical idea is to condition the policy on a speed-derived deadline signal (Td − t) and to clip the dense angle-progress reward so that the policy cannot exceed a per-step progress corresponding to the target speed. Simulation experiments evaluate reward ablations, horizon choices, and the two candidate speed-conditioning signals; the chosen 'time' conditioning with a randomized horizon offset produces good speed matching up to about 1.5 rad/s and saturation near 2.0 rad/s. Real-robot trials on the DLR-Hand II show robust reorientation at target speeds of 0.5, 1.0, and 2.0 rad/s, with 29/30 trials successful, but the reported execution times are determined by human inspection of video footage.","tokens_in":9190,"tokens_out":4701,"duration_ms":50747,"significance":"If the results hold, the paper offers a simple and effective way to make tactile in-hand manipulation policies speed-adjustable, and it reports faster reorientation than prior tactile-only work. The simulation evaluation is a clear strength: 1200 episodes per policy, multiple reward ablations, three training seeds, and a scatter plot that allows the speed-matching claim to be inspected directly. The zero-shot transfer to hardware is impressive qualitatively, but the quantitative speed-matching claim on the real robot rests on a much weaker measurement protocol than the simulation study. The 'fastest ever' claim is also not backed by a controlled comparison. The central idea is defensible, but the real-hardware evidence and some methodological details need to be strengthened or the claims need to be tempered.","major_comments":[{"comment":"The real-robot speed-matching evidence is not sufficient to support the abstract and contribution claims. Section V-B states that 'the success of the episode and the time required to reach the goal are determined by examining the video footage' and that 'for the few trials we checked carefully, we found that the required time closely matched the target speed (<10% off).' At ωd = 2.0 rad/s and a π rotation, the target time is about 1.6 s, so a 10% error is about 160 ms; visually judging when the cube crosses the 0.4 rad threshold from video has uncertainty on this order. With only 30 trials and no objective pose tracking, the zero-shot transfer of precise speed control is unquantified. Please either add an external pose-tracking measurement (or frame-by-frame video annotation with reported uncertainty) for all trials, or restrict the real-robot claims to qualitative robustness rather than speed matching.","section":"V-B"},{"comment":"The central mechanism for speed control, the clipping threshold θclip, is never defined as a function of the desired speed ωd. The text says the target speed 'is also used for the angle clipping operation,' but the formula is missing. Since θclip directly determines the maximum per-step angle progress, the paper must state how θclip is computed from ωd and the interaction frequency fnn (e.g., θclip = ωd / fnn). Without this, the method is not reproducible and the reader cannot judge whether the chosen clipping is consistent with the desired speed.","section":"V, Eq. (8)"},{"comment":"The 'fastest ever' claim is not supported by a controlled comparison. The comparison mixes different objects, tasks, and measurement protocols: Morgan et al. [5] use finger gaiting with visual pose tracking, Chen et al. [11] use diverse objects and an external camera, and Pitz et al. [16] report times for different objects and goals. Since the present evaluation is limited to cuboids with π/4-discretized goals and timing by video inspection, the claim that this is 'the fastest complex in-hand manipulation task that was ever shown on a real robot without vision' is overstated. Please either provide a matched comparison on the same objects and goal sets, or clearly state that the claim applies to this specific cuboid protocol and should not be read as a general benchmark.","section":"I-B / V-B"}],"minor_comments":[{"comment":"The sentence 'the magnitude gain for the respective network interaction frequency stays comparable' is unclear; please specify what 'magnitude gain' refers to and how it was measured.","section":"II-B"},{"comment":"The caption states that the lines are smoothed, but the smoothing procedure is not described; please specify the smoothing filter and window.","section":"Fig. 3"},{"comment":"The text says 'we did validate all three full-π rotations and two complex rotations' and that there were 30 trials, but the number of repetitions per condition is not reported; please give the trial counts for each goal and target speed.","section":"V-B"},{"comment":"The success threshold Δg = 0.4 rad is used in simulation, while real-robot goals are π/4-discretized; please clarify how the two thresholds are related when comparing real and simulated success rates.","section":"II-E"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about the real-robot timing limitation in V-B, but the abstract and contribution list make claims that are stronger than what that section can support. The simulation evidence is good, and the speed-adjustable idea is likely to be of interest to the robotics community. I would like the authors to either bring the real-robot measurement to the standard of the simulation evaluation or revise the claims accordingly. The missing definition of θclip is a concrete reproducibility issue that should be fixed. No code release is mentioned, which would be helpful given the sensitivity of the results to hyperparameters such as θclip and Hexp."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, this one deserves a real referee. The genuinely new thing is conditioning the policy on remaining time ξ = Td − t with a clipped angle-progress reward. It gives speed-adjustable SO(3) reorientation purely from tactile sensing in simulation, tracking target speeds well up to about 1.5 rad/s and saturating near 2.0 rad/s. That is a clean, useful result, and the simulation work is solid: 1200 episodes per policy, multiple reward ablations, three seeds. The authors are also honest about what didn't work—sparse time-optimal reward alone was inferior, and conditioning directly on target speed failed. That kind of negative result is informative.\n\nThe soft spot is the real-hardware part. The speed-matching claim and the 'fastest ever' headline rest on 30 trials where success and time were judged by a human from video. The authors say they couldn't run enough trials to compare quantitatively and that only a few trials were checked carefully, with '<10% off.' At 2 rad/s that is about 160 ms on a π rotation. That is not enough to quantify speed tracking on hardware or to support a comparative 'fastest ever' claim. The 29/30 success rate is good evidence of robust transfer, but speed accuracy on hardware is essentially unquantified. The stress-test note lands. This is fixable: objective pose tracking and more trials would settle it.\n\nMinor: no code or data release, which limits reproducibility. The 'time-optimal' framing is a bit loose since they don't prove optimality and the control frequency is a chosen hyperparameter, but they never claim global optimality—it's an objective for training, and they analyze it carefully.\n\nBottom line: the simulation result is a meaningful advance in the subfield, the reward design is well analyzed, and the hardware demo is suggestive. The paper should go to peer review with a request for stricter hardware evaluation and artifact release. It is not a desk reject. I would bring it to a robotics reading group and would cite it for the remaining-time conditioning trick.","headline":"Clever remaining-time conditioning gives a solid simulation result for speed-adjustable tactile in-hand reorientation, but the real-hardware speed claims rest on subjective video timing of 30 trials and outrun the evidence.","tokens_in":9781,"tokens_out":2638,"would_cite":true,"duration_ms":24129,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A torque-controlled robot hand can learn to reorient a grasped object to any SO(3) goal at a user-selected average speed using only tactile sensing, and the policy transfers to the real DLR-Hand II.","keywords":["tactile in-hand manipulation","reinforcement learning","speed-adjustable policies","time-optimal control","sim2real transfer","SO(3) reorientation","DLR-Hand II","clipped angle-progress reward"],"falsifier":"Run the published policy on the real DLR-Hand II with an external motion-capture or vision pose tracker and measure the time to reach the goal threshold across at least 30 trials at target speeds 0.5, 1.0, and 2.0 rad/s; if the measured effective speeds are consistently below the targets or the success rate falls well below the video-inspection estimate, the speed-control claim would be refuted.","tokens_in":8692,"feed_emoji":"🤖","tokens_out":5517,"duration_ms":48897,"temperature":0.7,"pith_summary":"The paper claims that a multi-fingered robot hand can learn to reorient a grasped cuboid to arbitrary SO(3) goals at a user-specified average speed, using only tactile sensing from joint angles and torques, and can do this faster than prior tactile-only approaches. The key is to condition the policy on the remaining time to a speed-derived deadline and to clip the angle-progress reward so that beating the target pace earns nothing extra. If correct, this gives robots a practical speed dial for in-hand manipulation without vision, which could be used to deliberately slow an assembly line to reduce wear. The paper also reports the fastest complex in-hand reorientation without visual input demonstrated on a real robot, with 29 of 30 real-world trials succeeding at target speeds of 0.5, 1.0, and 2.0 rad/s.","feed_headline":"Touch-only robot hand reorients a cube at your chosen speed","feed_subtitle":"Tactile RL policies trained on remaining time match target speeds up to 1.5 rad/s and transfer to the real hand.","key_machinery":"The load-bearing mechanism is the pair of a clipped angle-progress reward $r^{\\text{CL}}_t = \\lambda_\\theta \\min(\\theta_{t-1}-\\theta_t, \\theta_{\\text{clip}}) + r^{\\text{HE}}_t$ and the conditioning signal $\\xi = T_d - t$, where $T_d = \\theta_0/\\omega_d$ is the target time derived from the initial angular deviation and the requested speed. The clip removes any incentive to rotate faster than the pace, while the remaining-time signal gives the policy a countdown it can use to budget its finger gaits. Keeping the horizon partially observable (sampling an extra $H_{\\text{exp}}$ between 0 and 1 s) preserves the \"fear of termination\" that drives the policy to use the available time. Estimator-coupled reinforcement learning, in which the policy is trained together with the learned tactile state estimator under domain randomization, is what makes the speed control robust enough to transfer to the real hand.","core_discovery":"The paper establishes that a goal-conditioned SO(3) reorientation policy for a torque-controlled four-fingered hand can be made speed-adjustable by conditioning on the remaining time to a deadline rather than on the target speed itself. With a reward whose angle-progress term is clipped, so that rotating faster than the requested pace yields no extra reward, and with the episode horizon randomized to keep time partially observable, the policy learns to spend almost exactly the allotted time: in simulation it matches requested speeds up to about 1.5 rad/s and saturates near 2.0 rad/s, close to the estimated hardware ceiling of about 2.5 rad/s. The same policy, trained jointly with a tactile state estimator, transfers zero-shot to the real hand, succeeding in 29 of 30 video-verified trials at 0.5, 1.0, and 2.0 rad/s, which the authors present as the fastest complex in-hand reorientation without visual input demonstrated on a real robot.","pith_inferences":["Extension beyond cuboids is open: the paper trains and evaluates only cuboids with aspect ratios up to two, and the authors expect the method to hold for other objects, but a direct test on a sphere or irregular object is an untried experimental step.","The success of $\\xi = T_d - t$ over the raw target speed suggests the policy learns temporal budgeting rather than a static speed mapping, which is a testable hypothesis for other RL tasks with time-varying goals.","Objective real-world verification is still needed: because the real-robot timing relies on human video inspection, a pose-tracked repeat of the 30 trials would quantify the true distribution of reorientation times and success at 2.0 rad/s.","The saturation near 2.0 rad/s, below the oracle speed of about 2.5 rad/s, indicates that further speed gains are likely a matter of actuation and control frequency rather than reward design."],"forward_implications":["A user can set the average reorientation speed at deployment without retraining, across roughly 0.25 to 1.5 rad/s in simulation, with saturation near 2.0 rad/s.","The same reward-and-conditioning recipe, a clipped progress reward plus a remaining-time observation, may be portable to other partially observable RL control tasks that need adjustable tempo, such as legged locomotion.","Because only joint angles and torques are used, the hand can keep reorienting objects blindfolded, in environments where cameras are absent or unreliable.","In simulation the speed-adjustable agent maintains success above 93 percent, while the time-optimal variants exceed 95 percent, showing that speed control does not come at the cost of reliability."],"supporting_citations":[{"why":"Supplies the DLR-Hand II hardware platform used for the real-world validation.","marker":"[1]"},{"why":"Established the purely tactile in-hand manipulation setting with a torque-controlled hand that this work builds on.","marker":"[8]"},{"why":"Provided the modular reinforcement learning architecture and the π/4-discretized goals used in the real experiments.","marker":"[9]"},{"why":"Introduced estimator-coupled reinforcement learning, the training scheme that makes the policy robust to estimator inaccuracies and enables zero-shot transfer.","marker":"[10]"},{"why":"The prior shape-conditioned tactile agent this work extends and compares against for reorientation speed.","marker":"[16]"},{"why":"PPO is the reinforcement learning algorithm used to optimize the policies.","marker":"[20]"},{"why":"The GPU-accelerated simulator used for training with domain randomization.","marker":"[21]"}],"fun_headline_variants":["Touch-only robot hand reorients cube at your chosen speed","Speed-adjustable tactile in-hand manipulation via time-conditioned RL","RL policy for in-hand rotation: set pace, no camera needed","Time-optimal tactile reorientation transfers from sim to real hand"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The real-hardware timing and success are judged by a human operator watching video footage to see when the cube enters the goal threshold, and only 30 trials were run, so the headline speed-control result on hardware rests on sparse, subjective measurement rather than an external pose tracker.","fun_headline_variants_meta":{"raw":{"variants":["Touch-only robot hand reorients cube at your chosen speed","Speed-adjustable tactile in-hand manipulation via time-conditioned RL","RL policy for in-hand rotation: set pace, no camera needed","Time-optimal tactile reorientation transfers from sim to real hand"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000883,"raw_usage":{"total_tokens":3816,"prompt_tokens":947,"completion_tokens":2869,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":2796}},"tokens_in":563,"tokens_out":2869,"duration_ms":23537,"temperature":1.0,"reasoning_tokens":2796,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:47:05.737807+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the published policy on the real DLR-Hand II with an external motion-capture or vision pose tracker and measure the time to reach the goal threshold across at least 30 trials at target speeds 0.5, 1.0, and 2.0 rad/s; if the measured effective speeds are consistently below the targets or the success rate falls well below the video-inspection estimate, the speed-control claim would be refuted.","supporting_citations":[{"cited_title":"DLR-Hand II: Next generation of a dextrous robot hand,","cited_arxiv_id":null,"evidence_quote":"Supplies the DLR-Hand II hardware platform used for the real-world validation."},{"cited_title":"Learning purely tactile in-hand ma- nipulation with a torque-controlled hand,","cited_arxiv_id":null,"evidence_quote":"Established the purely tactile in-hand manipulation setting with a torque-controlled hand that this work builds on."},{"cited_title":"Dextrous tactile In-Hand manipulation using a modular reinforcement learning architecture,","cited_arxiv_id":null,"evidence_quote":"Provided the modular reinforcement learning architecture and the π/4-discretized goals used in the real experiments."},{"cited_title":"Estimator-Coupled reinforcement learning for robust purely tactile In-Hand manipulation,","cited_arxiv_id":null,"evidence_quote":"Introduced estimator-coupled reinforcement learning, the training scheme that makes the policy robust to estimator inaccuracies and enables zero-shot transfer."},{"cited_title":"Learning a shape- conditioned agent for purely tactile in-hand manipulation of various objects,","cited_arxiv_id":null,"evidence_quote":"The prior shape-conditioned tactile agent this work extends and compares against for reorientation speed."},{"cited_title":"Omniverse isaac gym reinforcement learning envi- ronments for isaac sim,","cited_arxiv_id":null,"evidence_quote":"The GPU-accelerated simulator used for training with domain randomization."}],"review_version":1}