Pith. sign in

REVIEW 9 cited by

Lessons from Learning to Spin "Pens"

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.18902 v2 pith:IKI67XGS submitted 2024-07-26 cs.RO cs.AIcs.LG

Lessons from Learning to Spin "Pens"

classification cs.RO cs.AIcs.LG
keywords policyobjectspen-likerealsimulationworldin-handlearning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In-hand manipulation of pen-like objects is an important skill in our daily lives, as many tools such as hammers and screwdrivers are similarly shaped. However, current learning-based methods struggle with this task due to a lack of high-quality demonstrations and the significant gap between simulation and the real world. In this work, we push the boundaries of learning-based in-hand manipulation systems by demonstrating the capability to spin pen-like objects. We first use reinforcement learning to train an oracle policy with privileged information and generate a high-fidelity trajectory dataset in simulation. This serves two purposes: 1) pre-training a sensorimotor policy in simulation; 2) conducting open-loop trajectory replay in the real world. We then fine-tune the sensorimotor policy using these real-world trajectories to adapt it to the real world dynamics. With less than 50 trajectories, our policy learns to rotate more than ten pen-like objects with different physical properties for multiple revolutions. We present a comprehensive analysis of our design choices and share the lessons learned during development.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SpikeATac: A Multimodal Tactile Finger with Taxelized Dynamic Sensing for Dexterous Manipulation

    cs.RO 2025-10 conditional novelty 7.0

    A fingertip with 16-taxel PVDF dynamic sensing plus capacitive static sensing enables fast delicate grasping and, with RLHF fine-tuning, in-hand manipulation of fragile objects.

  2. Learning to Play Piano in the Real World

    cs.RO 2025-03 unverdicted novelty 7.0

    A Sim2Real2Sim learning pipeline enables a real-world dexterous robot to play piano pieces including Happy Birthday and Ode to Joy with an average F1-score of 0.881.

  3. Towards Human-level Dexterous Teleoperation

    cs.RO 2026-07 conditional novelty 6.0

    A single-stage RL co-tracking controller trained on consecutive human-derived hand–object subgoals achieves ~75% real-robot success on long-horizon dexterous teleoperation where baselines fail.

  4. ETac: A Lightweight and Efficient Tactile Simulation Framework for Learning Dexterous Manipulation

    cs.RO 2026-04 unverdicted novelty 6.0

    ETac is a data-driven tactile simulation framework that matches FEM deformation accuracy at high speed, supporting 4096 parallel environments at 869 FPS and yielding 84.45% success in blind grasping across four object types.

  5. FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control

    cs.LG 2026-04 unverdicted novelty 6.0

    FlashSAC scales up Soft Actor-Critic with fewer updates, larger models, higher data throughput, and norm bounds to deliver faster, more stable training than PPO on high-dimensional robot control tasks across dozens of...

  6. FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control

    cs.LG 2026-04 unverdicted novelty 6.0

    FlashSAC improves training speed and final performance of off-policy RL on high-dimensional robot tasks by reducing update frequency, increasing model scale, and bounding norms to limit critic error accumulation.

  7. Grasp to Act: Dexterous Grasping for Tool Use in Dynamic Settings

    cs.RO 2026-02 conditional novelty 6.0

    Combining wrench-tested grasp optimization with real-time RL finger adjustments lets a 16-DoF robot hand keep tools stable during hammering, sawing, cutting, stirring, and scooping.

  8. One Hand to Rule Them All: Canonical Representations for Unified Dexterous Manipulation

    cs.RO 2026-02 unverdicted novelty 6.0

    A unified parameter space and canonical URDF enable cross-embodiment dexterous grasping policies with 81.9% zero-shot success on unseen hands like the 3-finger LEAP Hand.

  9. TopoRetarget: Interaction-Preserving Retargeting for Dexterous Manipulation

    cs.RO 2026-06 unverdicted novelty 5.0

    TopoRetarget uses a sparse interaction graph and distance-weighted Laplacian deformation optimization with kinematic and penetration constraints to retarget human demonstrations to dexterous hands while preserving tas...