REVIEW 3 cited by
LLM-Craft: Robotic Crafting of Elasto-Plastic Objects with Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
When humans create sculptures, we are able to reason about how geometrically we need to alter the clay state to reach our target goal. We are not computing point-wise similarity metrics, or reasoning about low-level positioning of our tools, but instead determining the higher-level changes that need to be made. In this work, we propose LLM-Craft, a novel pipeline that leverages large language models (LLMs) to iteratively reason about and generate deformation-based crafting action sequences. We simplify and couple the state and action representations to further encourage shape-based reasoning. To the best of our knowledge, LLM-Craft is the first system successfully leveraging LLMs for complex deformable object interactions. Through our experiments, we demonstrate that with the LLM-Craft framework, LLMs are able to successfully create a set of simple letter shapes. We explore a variety of rollout strategies, and compare performances of LLM-Craft variants with and without an explicit goal shape images. For videos and prompting details, please visit our project website: https://sites.google.com/andrew.cmu.edu/llmcraft/home
Forward citations
Cited by 3 Pith papers
-
LLM Trainer: Automated Robotic Data Generation via Demonstration Augmentation using LLMs
An LLM-based pipeline automatically augments one human demonstration into a large imitation-learning dataset, using Thompson sampling to pick the best annotation and beating expert-annotated baselines on most tasks.
-
PinchBot: Long-Horizon Deformable Manipulation with Guided Diffusion Policy
A single goal-conditioned diffusion policy, combined with pre-trained point cloud embeddings and collision-constrained action projection, can create pottery bowls of 8, 10, and 12 centimeter diameters.
-
Understanding Physical Properties of Unseen Deformable Objects by Leveraging Large Language Models and Robot Actions
Using robot actions and LLM visual reasoning, the system identifies deformability properties of unseen objects with up to 78.57% accuracy, which helps plan bin-packing at over 96% success after replanning.
Discussion (0). Sign in to comment.