REVIEW 4 major objections 4 minor 2 cited by
BiFold: Bimanual Cloth Folding with Language Guidance
T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read BiFold repurposes a pre-trained vision-language model to convert text commands into bimanual pick-and-place actions for cloth folding, reporting state-of-the-art results on an existing language-conditioned folding benchmark and the best…
desk verdict Solid empirical contribution with a useful auto-annotated bimanual dataset, but the language-guidance claim is untested because no ablation removes the language input. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing component is a contrastive vision-language transformer whose image and text branches are kept largely frozen and adapted with low-rank updates, followed by a transformer encoder that fuses token sequences and convolutional decoders that emit per-arm pick-and-place heatmaps. The other essential mechanism is the automatic dataset-annotation pipeline: it maps garment vertices to a per-category canonical coordinate space, thresholds those coordinates into semantic regions such as sleeves and waistbands, merges the left and right hand labels with a hand-designed rule table, and instantiates template sentences into hundreds of varied instructions. Together these allow the model to be trained on bimanual human demonstrations with no manual annotation.
What would settle it
Take a random sample of the new bimanual dataset, have a human annotator verify each automatically generated instruction and pick-and-place pair, and measure the label agreement rate; if a substantial fraction of the "place" labels or instructions are wrong, the bimanual evaluation is not a reliable measure of language-conditioned folding.
Extended reading notes
Core claim
The central discovery is that a frozen, low-rank-adapted vision-language transformer provides a sufficiently rich shared representation of garment images and natural-language folding instructions that the model can predict bimanual actions it was never explicitly taught. The policy produces pixel-space distributions over pick and place locations for each arm, constrained so that picks fall on the cloth mask, and it conditions on up to three previous keyframes to resolve ambiguities such as which side of a symmetric cloth is "top". The authors show that this design outperforms a prior transformer-based language-conditioned folding policy on an existing benchmark, and that on their own bimanual dataset it achieves the best average precision, lowest keypoint error, and lowest simulation mesh error while generalizing to new garments, new paraphrased instructions, and real images.
Load-bearing premise
The bimanual results rest on the automatic annotation pipeline producing correct language instructions and pick-and-place labels; if those labels are systematically noisy, the reported bimanual improvements would be overstated.
Editorial extensions
If this is right
- If the architecture is right, a frozen vision-language backbone plus small adaptation is enough for language-conditioned deformable-object manipulation, so the main barrier becomes labelled data rather than representation learning.
- The automatic annotation pipeline can be reapplied to other tracked demonstration datasets, lowering the cost of producing language-aligned manipulation benchmarks.
- Conditioning on a short history of keyframes improves pick-and-place precision in the bimanual setting, so memory of past states should be part of future folding policies.
- Predicting actions in pixel space, instead of on a downsampled point cloud, allows place positions to lie outside the current cloth silhouette, which the paper shows is needed for most bimanual folds.
- The reported gains on unseen tasks in the unimanual benchmark indicate that the text-image alignment transfers beyond the exact instruction templates used in training.
Reading between the lines
- A natural next step, which the paper leaves open, would be to couple BiFold's action heatmaps with an instruction-breaking planner so a single high-level goal yields a whole folding sequence; nothing in the paper rules this out.
- The annotation pipeline's reliance on canonical-coordinate thresholds assumes consistent garment topology within a category; testing it on highly varied designer garments would reveal how far the approach scales.
- Because the real-world evaluation is offline and qualitative, the strongest testable extension would be a full closed-loop dual-arm deployment on the same garments, measuring physical fold success rather than heatmap accuracy.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces BiFold, a vision-language model for bimanual cloth folding from RGB images and natural language instructions. The model uses a frozen SigLIP encoder adapted with LoRA, fuses image and text tokens in a transformer, optionally conditions on H previous observations, and outputs pick and place heatmaps for left and right arms. To train it, the authors augment the VR-Folding dataset with automatically generated language instructions via NOCS-based semantic labeling and template prompts. They evaluate on the unimanual Deng et al. benchmark (Table I), on their new bimanual dataset in simulation and image space (Tables II-III), and on real images offline (Fig. 4c). The paper claims state-of-the-art unimanual performance, best bimanual performance, and strong generalization to new instructions, garments, and environments.
Significance. If the results hold, the paper makes useful contributions: a practical recipe for adapting a pretrained vision-language model to bimanual cloth-folding action prediction, a fully automatic pipeline for generating language-aligned action labels from existing human demonstrations, and a new bimanual benchmark. The context mechanism and the move from point-cloud-based to pixel-space prediction are sensible design choices, and the external unimanual benchmark provides a useful sanity check. However, the central attribution of performance to language guidance is not yet supported by the experiments, and the evaluation lacks statistical grounding and quantitative real-world validation; these gaps currently limit the strength of the stated contributions.
major comments (4)
- [IV-D, Table IV] The central claim that language guidance is what enables BiFold's performance is not supported because no experiment removes or corrupts the language input. Section III-A defines the policy as πθ(at | ℓt, ot, ...) and all reported models retain the SigLIP text branch; Table IV only swaps the text encoder (T5) or changes the fusion/decoder architecture, never training a vision-only variant or one with scrambled text. Since the instruction templates are generated from the same semantic labels used to define the task, text is highly redundant with the visual observation and action distribution, and a vision-only policy with the same context could plausibly match many of the seen-instruction and unseen-instruction scores. The 'unseen task' rows in Table I are the only place where language seems necessary, and there BiFold is strong (e.g., Corner 100.0 with 1000 demonstrations), but without a no-language control the 'language guidance' attribution in the abstract remains unestablished.
- [IV-B, IV-C, Tables I-III] No error bars, confidence intervals, or training seeds are reported for any of the quantitative results. Tables I-III present single numbers for success rates, AP, KP-MSE, mIoU, and success; differences of a few percentage points (e.g., Table I, Half with 1000 demonstrations: BiFold 69.3% vs. Deng et al. 74.0%; Table III, Skirt SuccessIoU≥80: BiFold 31.7% vs. BiFold w/o context 34.9%) are within the range one would expect from stochastic training on small datasets, so the claimed state-of-the-art and consistent-outperformance conclusions are not statistically grounded.
- [IV-C, Fig. 4c] The real-world evaluation is offline and purely qualitative: the paper states 'we perform an offline qualitative evaluation on test images' and shows predicted actions in Fig. 4c, with no physical folding, no success metric, and no comparison to baselines. The abstract's claim of 'strong generalization to new instructions, garments, and environments' is therefore only partially supported; the 'environments' part is not demonstrated quantitatively. At minimum, the claim should be softened or the evaluation supplemented with a quantitative real-world study.
- [III-B, Algorithm 1, Appendix IV-B] The bimanual benchmark is self-created with an automatic annotation pipeline whose outputs are not validated. Algorithm 1 contains an explicit comment that 'Place vertices may be wrong,' and Appendix IV-B acknowledges that NOCS thresholding may fail for garment categories with high shape diversity. If the automatically parsed semantic labels and templates are systematically noisy, the bimanual results in Tables II and III and the associated generalization claims could be inflated. The paper should report annotation quality (e.g., human agreement on a sample, or a manual audit of parsed actions) and show that the reported bimanual results are not an artifact of label noise.
minor comments (4)
- [Appendix I-C] There is a typo in the sentence 'Wwe can observe that when the simulator becomes unstable...' — 'Wwe' should be 'We'.
- [Throughout] The reference to the prior work appears as 'Denget al.' without a space; it should be 'Deng et al.' in all occurrences.
- [IV-B, Table I] The baseline numbers for the unimanual benchmark are taken from Deng et al. without retraining; the paper should clarify whether the same data splits, augmentations, and evaluation protocol were used, since the comparison could be sensitive to such details.
- [III-A] The fixed context size H=3 is motivated by dataset statistics, but the paper does not discuss how the model behaves when a test sequence has more than three actions, which would require either truncating the context or using a longer horizon; a brief comment would clarify the expected failure mode.
Circularity Check
No significant circularity: the unimanual benchmark is external and the bimanual targets come from held-out human demonstrations, while the self-citations are not load-bearing.
full rationale
BiFold's derivation chain is self-contained and non-circular. The unimanual state-of-the-art claim (Table I) is evaluated on the external Deng et al. benchmark, with baseline numbers taken from that prior work and the model trained on that same fixed dataset; no parameter fitted by BiFold enters the definition of the target success metric. The bimanual claim is weaker because the dataset is self-created, but the ground-truth pick-and-place targets are human VR demonstration vertices re-rendered in simulation, not outputs of the model, and the test partition is held out. The language templates are generated from NOCS semantic labels, which may make the text branch partially redundant with the visual state, but this is a missing vision-only ablation rather than a circular derivation. The model is trained with a standard BCE loss against Gaussian-smoothed ground-truth positions, independent of the evaluation metrics. The paper explicitly flags its own limitations in Appendix IV-C (oracle ambiguity with human demonstrations, simulator physics, NOCS failure for diverse topologies) and in Section V, and these admissions do not conceal a circular step. Self-citations such as [46], [15], and [2] support only side claims about simulator inaccuracy, VR data collection, and bimanual benchmarking; they are not the load-bearing premise for the reported predictions. The strongest non-circular concern is the absence of a no-language or text-scrambled control, which affects the attribution of gains to language guidance, but that is an experimental-control issue, not a definitional or self-citation circularity.
Assumptions & free parameters
free parameters (6)
- Context window size H =
3
- Heatmap label variance Sigma =
5.0 I
- LoRA configuration =
rank 8, alpha 32, dropout 0.01
- Action filtering thresholds =
span > 5 frames, distance >= 0.1 m
- Simulation divergence filter =
z-score ratio > 3.5 filtered
- Success thresholds =
0.0125 m vertex distance; IoU >= 80%
assumptions (6)
- domain assumption SigLIP's pretrained vision-language representations transfer to cloth folding after LoRA adaptation
- domain assumption Pick-and-place primitives are sufficient to express all folding strategies
- domain assumption NOCS coordinates remain semantically consistent throughout a manipulation sequence
- domain assumption SoftGym cloth physics is accurate enough for evaluating folding success
- domain assumption A segmentation mask of the cloth is available at inference
- domain assumption Human VR demonstrations plus predefined folding protocols are valid ground-truth folding behavior
Cite this review
Pith. "Pith review of BiFold: Bimanual Cloth Folding with Language Guidance." pith.science (2026). https://pith.science/paper/UBMQQEK6
@misc{pith2026250116458,
author = {Pith},
title = {Pith review of: BiFold: Bimanual Cloth Folding with Language Guidance},
year = {2026},
howpublished = {\url{https://pith.science/paper/UBMQQEK6}},
note = {Machine review of arXiv:2501.16458}
}
read the original abstract
Cloth folding is a complex task due to the inevitable self-occlusions of clothes, their complicated dynamics, and the disparate materials, geometries, and textures that garments can have. In this work, we learn folding actions conditioned on text commands. Translating high-level, abstract instructions into precise robotic actions requires sophisticated language understanding and manipulation capabilities. To do that, we leverage a pre-trained vision-language model and repurpose it to predict manipulation actions. Our model, BiFold, can take context into account and achieves state-of-the-art performance on an existing language-conditioned folding benchmark. To address the lack of annotated bimanual folding data, we introduce a novel dataset with automatically parsed actions and language-aligned instructions, enabling better learning of text-conditioned manipulation. BiFold attains the best performance on our dataset and demonstrates strong generalization to new instructions, garments, and environments.
Figures
Figures from the paper (24 more)
Forward citations
Cited by 2 Pith papers
-
Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception
A robot folds cloth from spoken language by decomposing instructions with GPT-4o and grounding each step with a SigLIP2-based pick-and-place perception module.
-
Beyond Static Perception: Integrating Temporal Context into VLMs for Cloth Folding
Keyframe-based temporal context and LoRA fine-tuning improve language-guided pick-and-place predictions in the BiFold cloth-folding model.
Reference graph
Works this paper leans on
-
[1]
Household cloth object set: Fostering benchmarking in deformable object manipulation,
I. Garcia-Camacho, J. Borr `as, B. Calli, A. Norton, and G. Aleny `a, “Household cloth object set: Fostering benchmarking in deformable object manipulation,”IEEE Robotics and Automation Letters, vol. 7, no. 3, pp. 5866–5873, 2022
work page 2022
-
[2]
Benchmarking bimanual cloth manipulation,
I. Garcia-Camacho, M. Lippi, M. C. Welle, H. Yin, R. Antonova, A. Varava, J. Borras, C. Torras, A. Marino, G. Aleny `a, and D. Kragic, “Benchmarking bimanual cloth manipulation,”IEEE Robotics and Automation Letters, vol. 5, no. 2, pp. 1111–1118, 2020
work page 2020
-
[3]
Folds- former: Learning Sequential Multi-Step Cloth Manipulation With Space-Time Attention,
K. Mo, C. Xia, X. Wang, Y . Deng, X. Gao, and B. Liang, “Folds- former: Learning Sequential Multi-Step Cloth Manipulation With Space-Time Attention,”IEEE Robotics and Automation Letters, vol. 8, no. 2, pp. 760–767, 2023
work page 2023
-
[4]
CLIPort: What and Where Pathways for Robotic Manipulation,
M. Shridhar, L. Manuelli, and D. Fox, “CLIPort: What and Where Pathways for Robotic Manipulation,” inCoRL, 2021
work page 2021
-
[5]
Learning Language- Conditioned Deformable Object Manipulation with Graph Dynamics,
Y . Deng, K. Mo, C. Xia, and X. Wang, “Learning Language- Conditioned Deformable Object Manipulation with Graph Dynamics,” inICRA, 2024
work page 2024
-
[6]
SpeedFolding: Learning Efficient Bimanual Folding of Garments,
Y . Avigal, L. Berscheid, T. Asfour, T. Kr ¨oger, and K. Goldberg, “SpeedFolding: Learning Efficient Bimanual Folding of Garments,” inIROS, 2022, pp. 1–8
work page 2022
-
[7]
U-net: Convolutional networks for biomedical image segmentation,
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” inMedical Image Com- puting and Computer-Assisted Intervention, N. Navab, J. Hornegger, W. M. Wells, and A. F. Frangi, Eds., 2015, pp. 234–241
work page 2015
-
[8]
Cloth Funnels: Canonicalized-Alignment for Multi-Purpose Garment Manipulation,
A. Canberk, C. Chi, H. Ha, B. Burchfiel, E. Cousineau, S. Feng, and S. Song, “Cloth Funnels: Canonicalized-Alignment for Multi-Purpose Garment Manipulation,” inICRA, 2022
work page 2022
Show all 62 references
-
[9]
Unifolding: Towards sample-efficient, scalable, and generalizable robotic garment folding,
H. Xue, Y . Li, W. Xu, H. Li, D. Zheng, and C. Lu, “Unifolding: Towards sample-efficient, scalable, and generalizable robotic garment folding,” inCoRL, 2023
2023
-
[10]
Learning Transferable Visual Models From Natural Language Super- vision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning Transferable Visual Models From Natural Language Super- vision,” inICML, 2021
2021
-
[11]
Transporter networks: Rearranging the visual world for robotic ma- nipulation,
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V . Sindhwani, and J. Lee, “Transporter networks: Rearranging the visual world for robotic ma- nipulation,” inCoRL, 2020
2020
-
[12]
Attention is All you Need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is All you Need,” inNeurIPS, vol. 30, 2017
2017
-
[13]
Learning Visible Connec- tivity Dynamics for Cloth Smoothing,
X. Lin, Y . Wang, Z. Huang, and D. Held, “Learning Visible Connec- tivity Dynamics for Cloth Smoothing,” inCoRL, 2021
2021
-
[14]
GarmentTracking: Category-Level Garment Pose Tracking,
H. Xue, W. Xu, J. Zhang, T. Tang, Y . Li, W. Du, R. Ye, and C. Lu, “GarmentTracking: Category-Level Garment Pose Tracking,” inCVPR, June 2023, pp. 21 233–21 242
2023
-
[15]
A virtual reality framework for fast dataset creation applied to cloth manipulation with automatic semantic labelling,
J. Borr `as, A. Boix-Granell, S. Foix, and C. Torras, “A virtual reality framework for fast dataset creation applied to cloth manipulation with automatic semantic labelling,” inICRA, 2023, pp. 11 605–11 611
2023
-
[16]
Sigmoid loss for language image pre-training,
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” inICCV, 2023, pp. 11 975–11 986
2023
-
[17]
LoRA: Low-rank adaptation of large language models,
E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” inInternational Conference on Learning Representations, 2022. [Online]. Available: https://openreview.net/forum?id=nZeVKeeFYf9
2022
-
[18]
Reproducible scaling laws for contrastive language-image learning,
M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, and J. Jitsev, “Reproducible scaling laws for contrastive language-image learning,” inCVPR, 2023, pp. 2818–2829
2023
-
[19]
Eva-clip: Improved training techniques for clip at scale,
Q. Sun, Y . Fang, L. Wu, X. Wang, and Y . Cao, “Eva-clip: Improved training techniques for clip at scale,” arXiv:2303.15389, 2023
2023 arXiv
-
[20]
DINOv2: Learning robust visual features without supervision,
M. Oquab, T. Darcet, T. Moutakanni, H. V . V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. HAZIZA, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Je- gou, J. Mairal, P. Labatu...
2024
-
[21]
Masked autoencoders are scalable vision learners,
K. He, X. Chen, S. Xie, Y . Li, P. Doll ´ar, and R. Girshick, “Masked autoencoders are scalable vision learners,” inCVPR, 2022, pp. 15 979– 15 988
2022
-
[22]
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs,
S. Tong, E. Brown, P. Wu, S. Woo, M. Middepogu, S. C. Akula, J. Yang, S. Yang, A. Iyer, X. Pan, A. Wang, R. Fergus, Y . LeCun, and S. Xie, “Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs,” arXiv:2406.16860, 2024
2024 arXiv
-
[23]
Probing the 3d awareness of visual foundation models,
M. El Banani, A. Raj, K.-K. Maninis, A. Kar, Y . Li, M. Rubinstein, D. Sun, L. Guibas, J. Johnson, and V . Jampani, “Probing the 3d awareness of visual foundation models,” inCVPR, 2024, pp. 21 795– 21 806
2024
-
[24]
PaliGemma: A versatile 3B VLM for transfer,
L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, T. Un- terthiner, D. Keysers, S. Koppula, F. Liu, A. Grycner, A. Gritsenko, N. Houlsby, M. Kumar, K. Rong, J. Eisenschlos, R. Kabra, M. Bauer, M. Bo ...
2024 arXiv
-
[25]
Pali-3 vision language models: Smaller, faster, stronger,
X. Chen, X. Wang, L. Beyer, A. Kolesnikov, J. Wu, P. V oigtlaender, B. Mustafa, S. Goodman, I. Alabdulmohsin, P. Padlewski, D. Salz, X. Xiong, D. Vlasic, F. Pavetic, K. Rong, T. Yu, D. Keysers, X. Zhai, and R. Soricut, “Pali-3 vision language models: Smaller, faster, stronger,...
-
[26]
Spoc: Imitating shortest paths in simulation enables effective navigation and manipulation in the real world,
K. Ehsani, T. Gupta, R. Hendrix, J. Salvador, L. Weihs, K.-H. Zeng, K. P. Singh, Y . Kim, W. Han, A. Herrasti, R. Krishna, D. Schwenk, E. VanderBilt, and A. Kembhavi, “Spoc: Imitating shortest paths in simulation enables effective navigation and manipulation in the real world,...
2024
-
[27]
4M: Massively multimodal masked modeling,
D. Mizrahi, R. Bachmann, O. F. Kar, T. Yeo, M. Gao, A. Dehghan, and A. Zamir, “4M: Massively multimodal masked modeling,” inAdvances in Neural Information Processing Systems, 2023
2023
-
[28]
4M-21: An any-to-any vision model for tens of tasks and modalities,
R. Bachmann, O. F. Kar, D. Mizrahi, A. Garjani, M. Gao, D. Griffiths, J. Hu, A. Dehghan, and A. Zamir, “4M-21: An any-to-any vision model for tens of tasks and modalities,” arXiv:2406.09406, 2024
2024 arXiv
-
[29]
Octo: An open-source generalist robot policy,
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y . Tan, L. Y . Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine, “Octo: An open-source generalist robot policy,” inProceedings of Robotics...
2024
-
[30]
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control,
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, X. Chen, K. Choro- manski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Her- zog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalash- nikov, Y . Kuang...
2023 arXiv
-
[31]
Align before fuse: Vision and language representation learning with momentum distillation,
J. Li, R. R. Selvaraju, A. D. Gotmare, S. Joty, C. Xiong, and S. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” inNeurIPS, 2021
2021
-
[32]
CLOTH3D: Clothed 3D Humans,
H. Bertiche, M. Madadi, and S. Escalera, “CLOTH3D: Clothed 3D Humans,” inECCV, 2020, pp. 344–359
2020
-
[33]
Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation,
H. Wang, S. Sridhar, J. Huang, J. Valentin, S. Song, and L. J. Guibas, “Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation,” inCVPR, June 2019
2019
-
[34]
Posescript: 3d human poses from natural language,
G. Delmas, P. Weinzaepfel, T. Lucas, F. Moreno-Noguer, and G. Ro- gez, “Posescript: 3d human poses from natural language,” inECCV, 2022, p. 346–362
2022
-
[35]
Adam: A Method for Stochastic Optimiza- tion,
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimiza- tion,” inICLR, 2015, pp. 1–15
2015
-
[36]
SoftGym: Benchmarking Deep Reinforcement Learning for Deformable Object Manipulation,
X. Lin, Y . Wang, J. Olkin, and D. Held, “SoftGym: Benchmarking Deep Reinforcement Learning for Deformable Object Manipulation,” inCoRL, 2021
2021
-
[37]
Segment anything,
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, P. Dollar, and R. Gir- shick, “Segment anything,” inICCV, 2023, pp. 4015–4026
2023
-
[38]
Gpt-fabric: Folding and smoothing fabric by leveraging pre-trained foundation models,
V . Raval, E. Zhao, H. Zhang, S. Nikolaidis, and D. Seita, “Gpt-fabric: Folding and smoothing fabric by leveraging pre-trained foundation models,”arXiv preprint arXiv:2406.09640, 2024
2024 arXiv
-
[39]
Fab- ricflownet: Bimanual cloth manipulation with a flow-based policy,
T. Weng, S. Bajracharya, Y . Wang, K. Agrawal, and D. Held, “Fab- ricflownet: Bimanual cloth manipulation with a flow-based policy,” in CoRL, 2021
2021
-
[40]
Deep Imitation Learning of Sequential Fabric Smoothing From an Algorithmic Supervisor,
D. Seita, A. Ganapathi, R. Hoque, M. Hwang, E. Cen, A. K. Tanwani, A. Balakrishna, B. Thananjeyan, J. Ichnowski, N. Jamali, K. Yamane, S. Iba, J. Canny, and K. Goldberg, “Deep Imitation Learning of Sequential Fabric Smoothing From an Algorithmic Supervisor,” in IEEE/RSJ Intern...
2020
-
[41]
Flingbot: The unreasonable effectiveness of dynamic manipulation for cloth unfolding,
H. Ha and S. Song, “Flingbot: The unreasonable effectiveness of dynamic manipulation for cloth unfolding,” inConference on Robotic Learning (CoRL), 2021
2021
-
[42]
Film: Visual reasoning with a general conditioning layer,
E. Perez, F. Strub, H. de Vries, V . Dumoulin, and A. C. Courville, “Film: Visual reasoning with a general conditioning layer,” inAAAI, 2018
2018
-
[43]
Exploring the limits of transfer learning with a unified text-to-text transformer,
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,”Journal of Machine Learning Research, vol. 21, no. 140, pp. 1–67, 2020
2020
-
[44]
RT-1: Robotics Transformer for Real-World Control at Scale,
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. Joshi, R. Julian, D. Kalash- nikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Ma...
2022 arXiv
-
[45]
V oxposer: Composable 3d value maps for robotic manipulation with language models,
W. Huang, C. Wang, R. Zhang, Y . Li, J. Wu, and L. Fei-Fei, “V oxposer: Composable 3d value maps for robotic manipulation with language models,” inCoRL, 2023. [Online]. Available: https://openreview.net/forum?id=9 8LF30mOC
2023
-
[46]
Benchmarking the Sim-to-Real Gap in Cloth Manipulation,
D. Blanco-Mulero, O. Barbany, G. Alcan, A. Colom ´e, C. Torras, and V . Kyrki, “Benchmarking the Sim-to-Real Gap in Cloth Manipulation,” IEEE Robotics and Automation Letters, 2024
2024
-
[47]
Diffusion policy: Visuomotor policy learning via action diffusion,
C. Chi, S. Feng, Y . Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” in Proceedings of Robotics: Science and Systems (RSS), 2023
2023
-
[48]
Hierarchical diffusion policy for kinematics-aware multi-task robotic manipulation,
X. Ma, S. Patidar, I. Haughton, and S. James, “Hierarchical diffusion policy for kinematics-aware multi-task robotic manipulation,”CVPR, 2024
2024
-
[49]
Distilled Feature Fields Enable Few-Shot Language-Guided Manip- ulation,
W. Shen, G. Yang, A. Yu, J. Wong, L. P. Kaelbling, and P. Isola, “Distilled Feature Fields Enable Few-Shot Language-Guided Manip- ulation,” inCoRL, 2023
2023
-
[50]
ClothesNet: An Information-Rich 3D Garment Model Repository with Simulated Clothes Environment,
B. Zhou, H. Zhou, T. Liang, Q. Yu, S. Zhao, Y . Zeng, J. Lv, S. Luo, Q. Wang, X. Yu, H. Chen, C. Lu, and L. Shao, “ClothesNet: An Information-Rich 3D Garment Model Repository with Simulated Clothes Environment,” inICCV, 2023
2023
-
[51]
BlenderProc2: A Procedural Pipeline for Photorealistic Rendering,
M. Denninger, D. Winkelbauer, M. Sundermeyer, W. Boerdijk, M. Knauer, K. H. Strobl, M. Humt, and R. Triebel, “BlenderProc2: A Procedural Pipeline for Photorealistic Rendering,”Journal of Open Source Software, vol. 8, no. 82, p. 4901, 2023
2023
-
[52]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” inICLR, 2021. [Online]. Available: http...
2021
-
[53]
End-to-end object detection with transformers,
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV, 2020, p. 213–229
2020
-
[54]
PyBullet, a Python module for physics simulation for games, robotics and machine learning,
E. Coumans and Y . Bai, “PyBullet, a Python module for physics simulation for games, robotics and machine learning,” http://pybullet. org, 2016–2021
2016
-
[55]
MuJoCo: A physics engine for model-based control,
E. Todorov, T. Erez, and Y . Tassa, “MuJoCo: A physics engine for model-based control,” inIEEE/RSJ International Conference on Intelligent Robots and Systems, 2012. BiFold: Bimanual Cloth Folding with Language Guidance Supplementary Material CONTENTS I INTRODUCTION 1 II RELATE...
2012
-
[56]
Set pick and place heights using the radius of the picker, regardless of the world coordinate of the vertex
-
[57]
2https://huggingface.co/docs/transformers/en/model_doc/siglip Fig
The picker is moved to the picking position but at a predefined height. 2https://huggingface.co/docs/transformers/en/model_doc/siglip Fig. 12:Patch artifact:When using transformer decoders, we observe patch artifacts in the heatmap predictions
-
[58]
The picker moves to the pick position and closes the gripper
-
[59]
The picker moves to the position in 2)
-
[60]
The picker is moved to the placing position but at a predefined height
-
[61]
The picker goes to the placing position and opens the gripper
-
[62]
Fold a T-shirt into a square
The picker moves to the place position at the same predefined height as in 2). All the movements are performed at a speed of 5 mm/action except steps (2) and (7), which we perform 100 times faster as they are supposed to not interact with the cloth. The bimanual primitive uses...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.