Pith. sign in

REVIEW 1 cited by

Reward Learning from Suboptimal Demonstrations with Applications in Surgical Electrocautery

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.07185 v2 pith:NR7S3LF3 submitted 2024-04-10 cs.RO cs.AIcs.LG

Reward Learning from Suboptimal Demonstrations with Applications in Surgical Electrocautery

classification cs.RO cs.AIcs.LG
keywords demonstrationslearningrewardfunctionmethodsuboptimalsurgicalelectrocautery
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Automating robotic surgery via learning from demonstration (LfD) techniques is extremely challenging. This is because surgical tasks often involve sequential decision-making processes with complex interactions of physical objects and have low tolerance for mistakes. Prior works assume that all demonstrations are fully observable and optimal, which might not be practical in the real world. This paper introduces a sample-efficient method that learns a robust reward function from a limited amount of ranked suboptimal demonstrations consisting of partial-view point cloud observations. The method then learns a policy by optimizing the learned reward function using reinforcement learning (RL). We show that using a learned reward function to obtain a policy is more robust than pure imitation learning. We apply our approach on a physical surgical electrocautery task and demonstrate that our method can perform well even when the provided demonstrations are suboptimal and the observations are high-dimensional point clouds. Code and videos available here: https://sites.google.com/view/lfdinelectrocautery

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Feedback Matters: Augmenting Autonomous Dissection with Visual and Topological Feedback

    cs.RO 2025-10 conditional novelty 6.0

    A stretch-based tissue connectivity estimator plus an exposure-maximizing controller and recovery planner raised autonomous dissection success on a da Vinci robot to 80%.