REVIEW 4 major objections 5 minor 9 cited by
RoboChemist's dual-loop VLM+VLA design lifts average success by 23.57% and compliance by 0.298 on chemistry lab tasks.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 20:04 UTC pith:6A7IIFI4
load-bearing objection Solid VLM+VLA integration for robotic chemistry with a genuinely useful visual-prompting contribution, but the headline success-rate claim compares apples to oranges because the closed-loop system gets retries and the baselines don't. the 4 major comments →
RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
RoboChemist's central claim is that a VLM can serve three roles at once: planner, visual prompt generator, and monitor. Feeding the VLA model an image annotated with the VLM's bounding boxes and keypoints closes the gap between semantic instructions and precise manipulation. The VLA model is fine-tuned with prompted reference images alongside the usual camera views and text instructions, and the outer loop re-executes a primitive until the VLM's monitor confirms success. Evaluated on seven primitives (grasp, heat, pour, stir, transfer solid, insert, press button) and on complete protocol-like experiments such as acid-base neutralization and flame tests, the full system attains success rates
What carries the argument
The load-bearing mechanism is instruction-aware visual prompting: the VLM converts a subtask description plus safety guidelines into explicit 2D annotations (bounding boxes for regions of interest, points for grasp or target locations) overlaid on the RGB image, and these annotated images are used both during VLA fine-tuning and at inference as an extra input channel. The second component is the outer closed loop: after each primitive, the VLM inspects the current image and returns a success verdict, triggering re-execution when the step is judged incomplete, so a sequence of discrete actions can emulate contingent behavior such as 'pour until colorless.'
Load-bearing premise
The whole cascade depends on the vision-language model reliably placing bounding boxes and keypoints on transparent, deformable, and cluttered labware; if a prompt points to the wrong container or wrong grasp point, the action model will execute confidently on the wrong target, and the closed loop will then verify the wrong thing.
What would settle it
Run a pouring task with three visually similar transparent beakers containing colorless liquids in a cluttered scene, annotate ground-truth target containers by hand, and count how often the VLM's generated bounding box matches the intended container. If, without retraining, the prompt's spatial accuracy is near chance, the reported success-rate advantage should not transfer to scenes outside the training distribution. Alternatively, ablate the outer loop entirely (no monitor retry) and compare success rates: if the gap between the full system and the no-monitor version shrinks to near zero, t
If this is right
- A single system can execute multi-step chemistry protocols composed from a small set of trained primitives, with no task-specific programming beyond a natural-language description of apparatus and reagents.
- The closed-loop monitor converts a fixed script into condition-based behavior: retrying a grasp until it lands at the compliant position, or re-pouring acid until an indicator changes color.
- Visual prompts generated by a grounded vision-language model avoid the need for depth reconstruction of transparent labware, which previously caused failures in transparent-object manipulation.
- Reported generalization means a robot trained on seven primitives can be repurposed to new reagents, containers, and entire experiments by changing only the task description given to the VLM.
- The compliance metric shows that procedural adherence (not just task completion) can be evaluated and trained, which is necessary for real laboratory safety.
Where Pith is reading between the lines
- Inference not in the paper: the reported gains may come disproportionately from the retry loop rather than the prompts; an ablation that removes only the monitor while keeping prompted images would isolate the causal contribution of each loop.
- Inference not in the paper: if prompt accuracy is the bottleneck, combining the VLM's 2D marks with depth or segmentation only where transparency defeats RGB is a natural next step for cluttered scenes.
- Inference not in the paper: the same dual-loop architecture could transfer to other long-horizon, safety-critical manipulation domains, such as surgical assistance or assembly, where a semantic monitor can judge step completion from images.
- Inference not in the paper: a concrete stress test would vary the number of visually similar transparent containers and measure the VLM's prompt IoU against human labels; if prompt spatial accuracy degrades faster than downstream success, that would pinpoint where the cascade's reliability limit lives.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes RoboChemist, a dual-loop robotic chemistry system that couples a VLM (Qwen2.5-VL) with a VLA policy (π0). The VLM acts as planner, as generator of image-space visual prompts (bounding boxes and keypoints), and as monitor that verifies subtask completion and triggers re-execution. The VLA is fine-tuned on 400 demonstrations per primitive task, with a mixture of successful and second-attempt trajectories. Experiments cover seven primitive tasks and five complete chemistry protocols, comparing against ACT, RDT, and π0, plus visual-prompt baselines ReKep and MOKA. The headline claim is a 23.57 percentage-point higher average success rate and a 0.298 higher compliance rate over prior VLA baselines, with additional generalization results for unseen objects and workflows.
Significance. If the quantitative claims are robust, RoboChemist is a useful step toward closed-loop, safety-aware laboratory automation: it addresses transparent/deformable labware without depth reconstruction, and it integrates semantic monitoring into a VLA loop. The paper has real strengths: real-robot evaluation on a diverse chemistry task suite, a w/o-CL ablation showing that visual prompting alone improves over π0, and qualitative generalization to reaction types not seen in training. However, the current evaluation does not yet rigorously support the headline quantitative claim. The main comparison conflates the outer-loop retry mechanism with policy quality, the training-data mixture appears to be selected after seeing evaluation results, and the absence of statistical uncertainty makes the reported margins difficult to interpret. These are fixable with additional experiments and reporting, so the contribution is defensible in principle but needs a major revision.
major comments (4)
- [§3.3, Table 2] The headline 23.57 pp SR gain is not a like-for-like comparison. RoboChemist w/ CL re-executes a failed primitive until the VLM monitor declares success, whereas ACT/RDT/π0 are evaluated as single-pass policies (the paper states 'the loop would end after a failed attempt'). No baseline is augmented with the same monitor/retry wrapper, and no attempt counts or timeouts are reported. The w/o CL row (avg 82.14 vs π0's 70) shows visual prompting alone helps, but it does not decompose how much of the remaining 11.43 pp comes from retries versus policy quality. Add at least a π0+monitor ablation and report retry statistics.
- [§4.1/A.6, Table 7] The 300/100 training-data mixture (Config 2) used in the main experiments was selected after inspecting Table 7's evaluation results. This is test-set-based model selection and can inflate the reported numbers. The paper must either use a held-out validation set for this choice or report all configurations' end-to-end performance (with visual prompting and closed loop) so the reader can assess selection bias.
- [§4.1, Tables 2–3] The evaluation has 20 trials per task and no error bars, confidence intervals, or significance tests. Several SR differences are within binomial noise (e.g., 80 vs 85 in Table 2; 18/20 vs 17/20 in Table 1). Report 95% CIs or exact binomial tests for at least the headline averages, and for the compliance-rate differences, to support the claimed margins.
- [§3.2/A.8] The method's success depends on Qwen2.5-VL placing bounding boxes and keypoints correctly on transparent/deformable labware, but prompt accuracy is never measured independently. A.8 reports 'Prompting 35%' of 20 failures without defining the criterion or denominator, and the monitor's false positive/negative rates are unknown. Provide a human-annotated accuracy metric for generated prompts and monitor decisions on at least a subset of trials.
minor comments (5)
- [A.3.1, task 3] The task-decomposition prompt says 'flame test of copper(II) hydroxide' while A.2 defines the task as a CuSO4 flame test; this should be corrected.
- [Table 4, π0 row] The 'Press the Button' CR is 0.363, inconsistent with 0.575 in Table 2.
- [A.4] 'Manganese(II) hydroxide' should be the intended catalyst/species; as written the species/equation mismatch is confusing. Also, 'breaker' is a typo for 'beaker' in A.3.1 item 5.
- [Figure 18] The figure lacks axis labels and a legend; the text mentions seven variations but the figure shows only six labels.
- [A.6/Table 7] The relationship between Config 2's 70% average and the w/o CL average of 82.14% in Table 2 is unexplained; clarify whether visual prompts are included and whether Table 7 uses the same trial set as Table 2.
Circularity Check
No significant circularity: central comparisons use external human rubric and controlled ablations.
full rationale
The paper's headline SR/CR gains are measured against external baselines (ACT, RDT, pi0) fine-tuned on the same data and scored with a fixed human rubric in Appendix A.1. The visual prompting contribution is isolated by the w/o CL vs w/ CL comparison, and the closed-loop contribution is isolated by the w/o CL ablation. The VLM serves as planner/prompt-generator/monitor, which creates a potential self-referential loop, but the SR/CR metrics are externally defined in A.1, so the system's success is not defined as its own monitor's verdict. The retry-loop difference between RoboChemist w/CL and single-pass baselines is a fairness/experimental-design issue (missing baseline-with-monitor), not a circular reduction to the paper's inputs. Self-citations to related work (e.g., [16], [17], [33], [52], [81]) are contextual and not load-bearing. Therefore no circular step can be exhibited with a specific reduction.
Axiom & Free-Parameter Ledger
free parameters (3)
- Compliance rubric weights =
0, 0.25, 0.5, 0.75, 1 per task
- Training data mixture (successful vs second-attempt) =
300/100
- Number of trials per task =
20
axioms (4)
- domain assumption Qwen2.5-VL reliably grounds grasp and target points on transparent and cluttered scenes
- domain assumption Fine-tuned pi0 VLA can condition on the prompted image as an additional input channel
- domain assumption Human-defined compliance rubric reflects procedural safety norms
- domain assumption The 7 primitives compose into the 5 complete tasks
Cite this review
Pith. "Pith review of RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation." pith.science (2026). https://pith.science/paper/6A7IIFI4
@misc{pith2026250908820,
author = {Pith},
title = {Pith review of: RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/6A7IIFI4}},
note = {Machine review of arXiv:2509.08820}
}
read the original abstract
Robotic chemists promise to both liberate human experts from repetitive tasks and accelerate scientific discovery, yet remain in their infancy. Chemical experiments involve long-horizon procedures over hazardous and deformable substances, where success requires not only task completion but also strict compliance with experimental norms. To address these challenges, we propose \textit{RoboChemist}, a dual-loop framework that integrates Vision-Language Models (VLMs) with Vision-Language-Action (VLA) models. Unlike prior VLM-based systems (e.g., VoxPoser, ReKep) that rely on depth perception and struggle with transparent labware, and existing VLA systems (e.g., RDT, pi0) that lack semantic-level feedback for complex tasks, our method leverages a VLM to serve as (1) a planner to decompose tasks into primitive actions, (2) a visual prompt generator to guide VLA models, and (3) a monitor to assess task success and regulatory compliance. Notably, we introduce a VLA interface that accepts image-based visual targets from the VLM, enabling precise, goal-conditioned control. Our system successfully executes both primitive actions and complete multi-step chemistry protocols. Results show 23.57% higher average success rate and a 0.298 average increase in compliance rate over state-of-the-art VLA baselines, while also demonstrating strong generalization to objects and tasks.
Figures
Forward citations
Cited by 9 Pith papers
-
LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models
LIBERO-Safety supplies a scalable benchmark, data-generation pipeline, and 19,664-demonstration dataset that exposes a generalization-safety tension in current VLA models where diverse training improves collision avoi...
-
Dexora: Open-source VLA for High-DoF Bimanual Dexterity
Dexora is the first open-source VLA system for dual-arm dual-hand high-DoF manipulation, trained on 100K simulated and 10K real teleoperated trajectories with a discriminator-weighted diffusion policy, achieving 66.7%...
-
AEGIS: Assay-Aware Protocol Validation and Runtime Monitoring for Open-Source Liquid Handling Robots
AEGIS combines a rule-guided LLM protocol validator with a PCA/VLM visual runtime monitor to catch silent liquid-handling failures on the Opentrons OT-2, reporting adjusted F1 0.97 and average precision 0.89 on small ...
-
LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models
Introduces LIBERO-Safety benchmark with parametric scenario generation and 19,664 collision-free demonstrations, then evaluates VLA models to reveal a generalization-safety tension.
-
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation
A protocol-driven multi-agent VLA system with visual verification and AugSmolVLA augmentation improves wet-lab robot execution over ACT, X-VLA, and SmolVLA on atomic, composite, and bimanual tasks.
-
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation
Presents BioProVLA-Agent, a protocol-driven VLA-enabled multi-agent system for embodied biological manipulation with visual state verification and AugSmolVLA augmentation for robustness in wet-lab conditions.
-
Long-Term Memory for VLA-based Agents in Open-World Task Execution
ChemBot adds dual-layer memory and future-state asynchronous inference to VLA models, enabling better long-horizon success in chemical lab automation on collaborative robots.
-
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation
BioProVLA-Agent integrates protocol parsing, visual state verification, and VLA-based execution in a closed-loop multi-agent framework with AugSmolVLA augmentation to improve robustness for biological lab tasks like t...
-
Long-Term Memory for VLA-based Agents in Open-World Task Execution
A dual-layer memory and progress-aware VLA system for long-horizon chemical lab automation reports higher success rates than monolithic VLA baselines on a UR3 robot.
Reference graph
Works this paper leans on
-
[1]
Burger, P
B. Burger, P. M. Maffettone, V . V . Gusev, C. M. Aitchison, Y . Bai, X. Wang, X. Li, B. M. Alston, B. Li, R. Clowes, et al. A mobile robotic chemist.Nature, 583(7815):237–241, 2020
2020
-
[2]
N. J. Szymanski, B. Rendy, Y . Fei, R. E. Kumar, T. He, D. Milsted, M. J. McDermott, M. Gal- lant, E. D. Cubuk, A. Merchant, et al. An autonomous laboratory for the accelerated synthesis of novel materials.Nature, 624(7990):86–91, 2023
2023
-
[3]
T. Dai, S. Vijayakrishnan, F. T. Szczypi ´nski, J.-F. Ayme, E. Simaei, T. Fellowes, R. Clowes, L. Kotopanov, C. E. Shields, Z. Zhou, et al. Autonomous mobile robots for exploratory syn- thetic chemistry.Nature, pages 1–8, 2024
2024
-
[4]
D. A. Boiko, R. MacKnight, B. Kline, and G. Gomes. Autonomous chemical research with large language models.Nature, 624(7992):570–578, 2023
2023
-
[5]
Steiner, J
S. Steiner, J. Wolf, S. Glatzel, A. Andreou, J. M. Granda, G. Keenan, T. Hinkley, G. Aragon- Camarasa, P. J. Kitson, D. Angelone, et al. Organic synthesis in a modular robotic system driven by a chemical programming language.Science, 363(6423):eaav2211, 2019
2019
-
[6]
S. H. M. Mehr, M. Craven, A. I. Leonov, G. Keenan, and L. Cronin. A universal system for digitization and automatic execution of the chemical synthesis literature.Science, 370(6512): 101–108, 2020
2020
-
[7]
C. W. Coley, D. A. Thomas III, J. A. Lummiss, J. N. Jaworski, C. P. Breen, V . Schultz, T. Hart, J. S. Fishman, L. Rogers, H. Gao, et al. A robotic platform for flow synthesis of organic compounds informed by ai planning.Science, 365(6453):eaax1566, 2019
2019
-
[8]
H. Liu, C. Li, Q. Wu, and Y . J. Lee. Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
2023
-
[9]
H. Liu, C. Li, Y . Li, and Y . J. Lee. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26296–26306, 2024
2024
-
[10]
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 24185–24198, 2024. 9
2024
-
[11]
S. Tong, E. Brown, P. Wu, S. Woo, M. Middepogu, S. C. Akula, J. Yang, S. Yang, A. Iyer, X. Pan, A. Wang, R. Fergus, Y . LeCun, and S. Xie. Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37: 87310–87356, 2024
2024
-
[12]
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al. Qwen2. 5-vl technical report.arXiv preprint arXiv:2502.13923, 2025
Pith/arXiv arXiv 2025
-
[13]
X. Liu, B. Tian, Z. Wang, R. Wang, K. Sheng, B. Zhang, H. Zhao, and G. Zhou. Delving into shape-aware zero-shot semantic segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2999–3009, 2023
2023
-
[14]
P. Li, B. Tian, Y . Shi, X. Chen, H. Zhao, G. Zhou, and Y .-Q. Zhang. Toist: Task oriented instance segmentation transformer with noun-pronoun distillation.Advances in Neural Infor- mation Processing Systems, 35:17597–17611, 2022
2022
-
[15]
B. Jin, Y . Zheng, P. Li, W. Li, Y . Zheng, S. Hu, X. Liu, J. Zhu, Z. Yan, H. Sun, et al. Tod3cap: Towards 3d dense captioning in outdoor scenes. InEuropean Conference on Computer Vision, pages 367–384. Springer, 2024
2024
-
[16]
H. Chi, H.-a. Gao, Z. Liu, J. Liu, C. Liu, J. Li, K. Yang, Y . Yu, Z. Wang, W. Li, et al. Impromptu vla: Open weights and open data for driving vision-language-action models.arXiv preprint arXiv:2505.23757, 2025
Pith/arXiv arXiv 2025
-
[17]
K. Ding, B. Chen, Y . Su, H.-a. Gao, B. Jin, C. Sima, W. Zhang, X. Li, P. Barsch, H. Li, et al. Hint-ad: Holistically aligned interpretability in end-to-end autonomous driving.arXiv preprint arXiv:2409.06702, 2024
Pith/arXiv arXiv 2024
-
[18]
M. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024
Pith/arXiv arXiv 2024
-
[19]
Ghosh, H
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y . Tan, L. Y . Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine. Octo: An open-source generalist robot policy. InProceedings of Robotics: Science and Systems, Delft, Netherlands, 2024
2024
-
[20]
S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu. Rdt-1b: a diffusion foundation model for bimanual manipulation.International Conference on Learning Representations, 2025
2025
-
[21]
Black, N
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Haus- man, B. Ichter, et al.π 0: A vision-language-action flow model for general robot control. In Conference on Robot Learning. PMLR, 2024
2024
-
[22]
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine. Fast: Efficient action tokenization for vision-language-action models.arXiv preprint arXiv:2501.09747, 2025
Pith/arXiv arXiv 2025
-
[23]
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, et al.π 0.5: a vision-language-action model with open-world generalization. arXiv preprint arXiv:2504.16054, 2025
Pith/arXiv arXiv 2025
-
[24]
W. Huang, C. Wang, R. Zhang, Y . Li, J. Wu, and L. Fei-Fei. V oxposer: Composable 3d value maps for robotic manipulation with language models.arXiv preprint arXiv:2307.05973, 2023
Pith/arXiv arXiv 2023
-
[25]
W. Huang, C. Wang, Y . Li, R. Zhang, and L. Fei-Fei. Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation.arXiv preprint arXiv:2409.01652, 2024. 10
Pith/arXiv arXiv 2024
-
[26]
Y . R. Wang, Y . Zhao, H. Xu, S. Eppel, A. Aspuru-Guzik, F. Shkurti, and A. Garg. Mv- trans: Multi-view perception of transparent objects. In2023 IEEE International Conference on Robotics and Automation (ICRA), pages 3771–3778. IEEE, 2023
2023
-
[27]
Brown, B
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020
1901
-
[28]
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023
Pith/arXiv arXiv 2023
-
[29]
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023
Pith/arXiv arXiv 2023
-
[30]
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288, 2023
Pith/arXiv arXiv 2023
-
[31]
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, et al. Qwen2. 5 technical report.arXiv preprint arXiv:2412.15115, 2024
Pith/arXiv arXiv 2024
-
[32]
X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Dong, Q. Du, Z. Fu, et al. Deepseek llm: Scaling open-source language models with longtermism.arXiv preprint arXiv:2401.02954, 2024
Pith/arXiv arXiv 2024
-
[33]
Z. Zhang, X. Li, S. Zou, G. Chi, S. Li, X. Qiu, G. Wang, G. Zheng, L. Wang, H. Zhao, et al. Chameleon: Fast-slow neuro-symbolic lane topology extraction.arXiv preprint arXiv:2503.07485, 2025
Pith/arXiv arXiv 2025
-
[34]
Zitkovich, T
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, Q. Vuong, V . Vanhoucke, H. Tran, R. Soricut, A. Singh, J. Singh, P. Sermanet, P. R. San- keti, G. Salazar, M. S. Ryoo, K. Reymann, K. Rao, K. Pertsch, I. Mordatch, H. Michalewski, Y . Lu, S. Levine, L. Lee, T.-W. E. Lee, I. Leal, Y . Kuang, D. Kalashnikov, R. Jul...
2023
-
[35]
Mandlekar, Y
A. Mandlekar, Y . Zhu, A. Garg, J. Booher, M. Spero, A. Tung, J. Gao, J. Emmons, A. Gupta, E. Orbay, et al. Roboturk: A crowdsourcing platform for robotic skill learning through imita- tion. InConference on Robot Learning, pages 879–893. PMLR, 2018
2018
-
[36]
Ebert, Y
F. Ebert, Y . Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine. Bridge data: Boosting generalization of robotic skills with cross-domain datasets. InRobotics: Science and Systems, New York City, USA, 2022
2022
-
[37]
O. X.-E. Collaboration, A. O’Neill, A. Rehman, A. Gupta, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, A. Tung, A. Bewley, A. Her- zog, A. Irpan, A. Khazatsky, A. Rai, A. Gupta, A. Wang, A. Kolobov, A. Singh, A. Garg, A. Kembhavi, A. Xie, A. Brohan, A. Raffin, A. Sharma, A. Yavary, A. Jain, A. Balakrishna, A. Wah...
2024
-
[38]
Khazatsky, K
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, P. D. Fagan, J. Hejna, M. Itkina, M. Lepert, Y . J. Ma, P. T. Miller, J. Wu, S. Belkhale, S. Dass, H. Ha, A. Jain, A. Lee, Y . Lee, M. Memmel, S. Park, I. Radosavovic, K. Wang, A. Zhan, K. Black, C. Chi, K. B. Hatch, S. Lin, J. ...
2024
-
[39]
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. Joshi, R. Ju- lian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, ...
Pith/arXiv arXiv 2022
-
[40]
AgiBot-World-Contributors, Q. Bu, J. Cai, L. Chen, X. Cui, Y . Ding, S. Feng, S. Gao, X. He, X. Hu, X. Huang, S. Jiang, Y . Jiang, C. Jing, H. Li, J. Li, C. Liu, Y . Liu, Y . Lu, J. Luo, P. Luo, Y . Mu, Y . Niu, Y . Pan, J. Pang, Y . Qiao, G. Ren, C. Ruan, J. Shan, Y . Shen, C. Shi, M. Shi, M. Shi, C. Sima, J. Song, H. Wang, W. Wang, D. Wei, C. Xie, G. Xu...
Pith/arXiv arXiv 2025
-
[41]
X. Li, M. Liu, H. Zhang, C. Yu, J. Xu, H. Wu, C. Cheang, Y . Jing, W. Zhang, H. Liu, et al. Vision-language foundation models as effective robot imitators.International Conference on Learning Representations, 2024
2024
-
[42]
Huang, S
J. Huang, S. Yong, X. Ma, X. Linghu, P. Li, Y . Wang, Q. Li, S.-C. Zhu, B. Jia, and S. Huang. An embodied generalist agent in 3D world. InProceedings of the 41st International Conference on Machine Learning, pages 20413–20451. PMLR, 2024
2024
-
[43]
Z. Durante, B. Sarkar, R. Gong, R. Taori, Y . Noda, P. Tang, E. Adeli, S. K. Lakshmikanth, K. Schulman, A. Milstein, et al. An interactive agent foundation model.arXiv preprint arXiv:2402.05929, 2024
Pith/arXiv arXiv 2024
-
[44]
R. Zheng, Y . Liang, S. Huang, J. Gao, H. Daum ´e III, A. Kolobov, F. Huang, and J. Yang. Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies.arXiv preprint arXiv:2412.10345, 2024
Pith/arXiv arXiv 2024
-
[45]
J. Zheng, J. Li, D. Liu, Y . Zheng, Z. Wang, Z. Ou, Y . Liu, J. Liu, Y .-Q. Zhang, and X. Zhan. Universal actions for enhanced embodied foundation models.arXiv preprint arXiv:2501.10105, 2025
Pith/arXiv arXiv 2025
-
[46]
X. Li, P. Li, M. Liu, D. Wang, J. Liu, B. Kang, X. Ma, T. Kong, H. Zhang, and H. Liu. Towards generalist robot policies: What matters in building vision-language-action models. arXiv preprint arXiv:2412.14058, 2024
Pith/arXiv arXiv 2024
-
[47]
J. Wen, Y . Zhu, J. Li, M. Zhu, Z. Tang, K. Wu, Z. Xu, N. Liu, R. Cheng, C. Shen, et al. Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation. IEEE Robotics and Automation Letters, 2025
2025
-
[48]
Z. Wu, Y . Zhou, X. Xu, Z. Wang, and H. Yan. Momanipvla: Transferring vision-language- action models for general mobile manipulation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025
2025
-
[49]
A. Jiang, Y . Gao, Z. Sun, Y . Wang, J. Wang, J. Chai, Q. Cao, Y . Heng, H. Jiang, Y . Dong, et al. Diffvla: Vision-language guided diffusion planning for autonomous driving.arXiv preprint arXiv:2505.19381, 2025
Pith/arXiv arXiv 2025
-
[50]
Q. Li, Y . Liang, Z. Wang, L. Luo, X. Chen, M. Liao, F. Wei, Y . Deng, S. Xu, Y . Zhang, et al. Cogact: A foundational vision-language-action model for synergizing cognition and action in robotic manipulation.arXiv preprint arXiv:2411.19650, 2024
Pith/arXiv arXiv 2024
-
[51]
D. Qu, H. Song, Q. Chen, Y . Yao, X. Ye, Y . Ding, Z. Wang, J. Gu, B. Zhao, D. Wang, et al. Spatialvla: Exploring spatial representations for visual-language-action model.arXiv preprint arXiv:2501.15830, 2025
Pith/arXiv arXiv 2025
-
[52]
K. Ding, B. Chen, R. Wu, Y . Li, Z. Zhang, H.-a. Gao, S. Li, G. Zhou, Y . Zhu, H. Dong, et al. Preafford: Universal affordance-based pre-grasping for diverse objects and environments. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 7278–7285. IEEE, 2024
2024
-
[53]
J. Liu, M. Liu, Z. Wang, P. An, X. Li, K. Zhou, S. Yang, R. Zhang, Y . Guo, and S. Zhang. Robomamba: Efficient vision-language-action model for robotic reasoning and manipulation. Advances in Neural Information Processing Systems, 37:40085–40110, 2024
2024
-
[54]
H. Zhen, X. Qiu, P. Chen, J. Yang, X. Yan, Y . Du, Y . Hong, and C. Gan. 3D-VLA: A 3D vision- language-action generative world model. InProceedings of the 41st International Conference on Machine Learning, pages 61229–61245. PMLR, 2024. 13
2024
-
[55]
J. Bjorck, F. Casta ˜neda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y . Fang, D. Fox, F. Hu, S. Huang, et al. Gr00t n1: An open foundation model for generalist humanoid robots.arXiv preprint arXiv:2503.14734, 2025
Pith/arXiv arXiv 2025
-
[56]
G. R. Team, S. Abeyruwan, J. Ainslie, J.-B. Alayrac, M. G. Arenas, T. Armstrong, A. Balakr- ishna, R. Baruch, M. Bauza, M. Blokzijl, et al. Gemini robotics: Bringing ai into the physical world.arXiv preprint arXiv:2503.20020, 2025
Pith/arXiv arXiv 2025
-
[57]
J. Liu, H. Chen, P. An, Z. Liu, R. Zhang, C. Gu, X. Li, Z. Guo, S. Chen, M. Liu, et al. Hy- bridvla: Collaborative diffusion and autoregression in a unified vision-language-action model. arXiv preprint arXiv:2503.10631, 2025
Pith/arXiv arXiv 2025
-
[58]
A. Bar, Y . Gandelsman, T. Darrell, A. Globerson, and A. Efros. Visual prompting via image inpainting.Advances in Neural Information Processing Systems, 35:25005–25017, 2022
2022
-
[59]
M. Jia, L. Tang, B.-C. Chen, C. Cardie, S. Belongie, B. Hariharan, and S.-N. Lim. Visual prompt tuning. InEuropean conference on computer vision, pages 709–727. Springer, 2022
2022
-
[60]
J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao. Set-of-mark prompting unleashes extraor- dinary visual grounding in gpt-4v.arXiv preprint arXiv:2310.11441, 2023
Pith/arXiv arXiv 2023
-
[61]
S. Yoo, E. Kim, D. Jung, J. Lee, and S. Yoon. Improving visual prompt tuning for self- supervised vision transformers. InInternational Conference on Machine Learning, pages 40075–40092. PMLR, 2023
2023
-
[62]
W. Liu, X. Shen, C.-M. Pun, and X. Cun. Explicit visual prompting for low-level structure segmentations. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19434–19445, 2023
2023
-
[63]
F. Li, Q. Jiang, H. Zhang, T. Ren, S. Liu, X. Zou, H. Xu, H. Li, J. Yang, C. Li, et al. Visual in-context prompting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12861–12871, 2024
2024
-
[64]
M. Cai, H. Liu, S. K. Mustikovela, G. P. Meyer, Y . Chai, D. Park, and Y . J. Lee. Vip-llava: Making large multimodal models understand arbitrary visual prompts. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12914–12923, 2024
2024
-
[65]
C. Xu, Y . Zhu, H. Shen, B. Chen, Y . Liao, X. Chen, and L. Wang. Progressive visual prompt learning with contrastive feature re-formation.International Journal of Computer Vision, 133 (2):511–526, 2025
2025
-
[66]
Kirillov, E
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, et al. Segment anything. InProceedings of the IEEE/CVF international conference on computer vision, pages 4015–4026, 2023
2023
-
[67]
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haz- iza, F. Massa, A. El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023
Pith/arXiv arXiv 2023
-
[68]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervi- sion. InInternational conference on machine learning, pages 8748–8763. PmLR, 2021
2021
-
[69]
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. De- hghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transform- ers for image recognition at scale.arXiv preprint arXiv:2010.11929, 2020
Pith/arXiv arXiv 2010
-
[70]
Moenning and N
C. Moenning and N. A. Dodgson. Fast marching farthest point sampling. Technical report, University of Cambridge, Computer Laboratory, 2003. 14
2003
-
[71]
Krishna and M
K. Krishna and M. N. Murty. Genetic k-means algorithm.IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 29(3):433–439, 1999
1999
-
[72]
Z. Liu, M. Zhang, and Y . Li. Kuda: Keypoints to unify dynamics learning and visual prompting for open-vocabulary robotic manipulation.arXiv preprint arXiv:2503.10546, 2025
Pith/arXiv arXiv 2025
-
[73]
K. Fang, F. Liu, P. Abbeel, and S. Levine. Moka: Open-world robotic manipulation through mark-based visual prompting.Robotics: Science and Systems (RSS), 2024
2024
-
[74]
Harazono, H
Y . Harazono, H. Shimono, K. Hata, T. Mitsuyama, and T. Horinouchi. Evaluation of microplate handling accuracy for applying robotic arms in laboratory automation.SLAS technology, 29 (6):100200, 2024
2024
-
[75]
N. Yoshikawa, A. Z. Li, K. Darvish, Y . Zhao, H. Xu, A. Kuramshin, A. Aspuru-Guzik, A. Garg, and F. Shkurti. Chemistry lab automation via constrained task and motion planning.arXiv preprint arXiv:2212.09672, 2022
Pith/arXiv arXiv 2022
-
[76]
Darvish, M
K. Darvish, M. Skreta, Y . Zhao, N. Yoshikawa, S. Som, M. Bogdanovic, Y . Cao, H. Hao, H. Xu, A. Aspuru-Guzik, et al. Organa: a robotic assistant for automated chemistry experimentation and characterization.Matter, 8(2), 2025
2025
-
[77]
Fakhruldeen, G
H. Fakhruldeen, G. Pizzuto, J. Glowacki, and A. I. Cooper. Archemist: Autonomous robotic chemistry system architecture. In2022 International Conference on Robotics and Automation (ICRA), pages 6013–6019. IEEE, 2022
2022
-
[78]
Knobbe, H
D. Knobbe, H. Zwirnmann, M. Eckhoff, and S. Haddadin. Core processes in intelligent robotic lab assistants: Flexible liquid handling. In2022 IEEE/RSJ international conference on intelli- gent robots and systems (IROS), pages 2335–2342. IEEE, 2022
2022
-
[79]
Schober, R
D. Schober, R. G ¨uldenring, J. Love, and L. Nalpantidis. Vision-based robot manipulation of transparent liquid containers in a laboratory setting. In2025 IEEE/SICE International Sympo- sium on System Integration (SII), pages 1193–1200. IEEE, 2025
2025
-
[80]
S. Li, Y . Huang, C. Guo, T. Wu, J. Zhang, L. Zhang, and W. Ding. Chemistry3d: Robotic interaction benchmark for chemistry experiments.arXiv preprint arXiv:2406.08160, 2024
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.