{"id":"7803670e-05f4-455c-b268-32fab36f39f3","arxiv_id":"2603.09971","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"TiPToP, a zero-training modular planner using pretrained vision-language models and GPU-accelerated TAMP, achieves 74.6% success over 165 trials versus 52.4% for the 350-hour-trained pi0.5-DROID baseline across 28 manipulation scenes.","lead":"TiPToP is a robot control system that combines off-the-shelf vision and language models with a GPU-powered task-and-motion planner, so it can follow natural-language instructions without any robot training data. In head-to-head tests it matched or beat a state-of-the-art policy trained on 350 hours of demonstrations, while planning and executing faster.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The π0.5 comparison uses unmatched termination limits — 800 steps externally vs 120 s internally — concentrated in the semantic scenes that drive TiPToP's margin; this needs a matched-timeout rerun before the headline success gap is accepted.","rationale":"The paper's central claim is that a zero-training modular planner can match or outperform a fine-tuned VLA. The evidence hinges on the aggregate success rate across 28 scenes. The largest gaps appear in semantic and multi-step categories, and most of those scenes were run by the external team with a 53 s wall-clock cap for π0.5. This asymmetry is a concrete, correctable flaw rather than a disagreement with the VLA consensus: it directly affects the numerical headline. The reader's weakest assumption ('evaluation set and protocol do not favor TiPToP') captures this, but I isolate the termination-limit asymmetry as the most actionable component. The proposed rerun with matched 120 s timeout would settle whether π0.5's poor semantic performance is intrinsic or partly an artifact of truncation. I do not see an internal inconsistency in the architecture claims; the open-loop limitation is honestly acknowledged and analyzed. The missing MolmoSpaces results are a report-quality issue, not the central comparison. Therefore the reader's CONDITIONAL verdict remains appropriate.","tokens_in":20975,"tokens_out":10953,"duration_ms":98290,"concrete_test":"Rerun the eight unmarked semantic scenes and AirPods→cup (45 trials) using the internal protocol: 120 s wall-clock timeout for π0.5, same initial configurations, 5 trials per scene, recording time-to-success and per-step logs. If π0.5's aggregate success on these scenes increases substantially (e.g., from 10/40 to ≥20/40, or any 0/5 scene becomes ≥3/5), the reported success-rate gap is inflated by the 800-step cap; if it stays at 10/40–13/40, the concern is refuted and the conditional acceptance should hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"TiPToP's headline advantage over π0.5-DROID rests heavily on the externally-run scenes (unmarked in Table I), where π0.5 scored 10/40 in the semantic category and 3/5 in AirPods→cup. Appendix C states that external evaluators terminated π0.5 after 800 control steps (~53 s at 15 Hz), whereas system-designer runs used a 120 s timeout. The paper asserts this limit is 'generous,' but reports no completion-time distribution for π0.5. Because π0.5 is a reactive policy that may need multiple grasp attempts and closed-loop recovery, a 53 s cap can convert partial progress into failures precisely on the semantic/multi-step tasks that generate most of TiPToP's aggregate margin (74.6% vs 52.4%). The central claim 'matches or outperforms' is therefore not established until π0.5 is given the same wall-clock budget in the external scenes. Additionally, the abstract's MolmoSpaces 'ranks first' claim appears nowhere in the body, so it cannot be independently weighed; however, the π0.5 comparison is the main issue.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents TiPToP, a modular open-vocabulary manipulation system that combines pretrained perception models (FoundationStereo, M2T2, SAM-2, Gemini Robotics-ER) with the GPU-parallelized task-and-motion planner cuTAMP. From a single stereo image pair and a natural-language instruction, TiPToP builds an object-centric scene representation, produces a symbolic goal, plans a full pick-and-place trajectory, and executes it open-loop with a joint impedance controller. The main empirical claim is that TiPToP, which requires no robot training data, 'matches or outperforms' π0.5-DROID, a VLA fine-tuned on 350 hours of DROID data, across 28 evaluation scenes: aggregate success 98/165 (74.6%) vs. 55/165 (52.4%), with faster average completion time on five of six measured scenes. The evaluation includes simulation, an in-house DROID setup, and an external DROID setup operated by a separate team, plus deployment on UR5e and WidowX and a wiping extension. A failure analysis over 173 additional trials attributes most failures to grasping, mesh approximation, VLM detection, and planner timeouts. The abstract also claims first place on the MolmoSpaces benchmark, though this is not described in the body.","tokens_in":21245,"tokens_out":8395,"duration_ms":78739,"significance":"If the claimed result holds, it is significant: a zero-robot-data modular system composed of off-the-shelf foundation models and geometric planning would be a competitive alternative to a VLA fine-tuned on embodiment-specific demonstrations, and the modular architecture provides a practical route for component-level debugging. The paper has real strengths: it releases open-source code, uses an external evaluation team for part of the study, reports both success rate and task progress, and honestly discusses open-loop execution as a key limitation. However, the headline comparison is not yet convincing as stated. The termination budget for π0.5-DROID differs between the external and designer-run scenes, real-world scenes use only five trials with no significance testing, and the task menu is author-selected. The conclusion 'matches or outperforms' therefore needs a matched-protocol rerun or equivalent evidence before it can be accepted.","major_comments":[{"comment":"The unmarked scenes in Table I were run by the external team, where π0.5-DROID trials were terminated after 800 control steps; at 15 Hz this is ≈53 s. The dagger-marked designer scenes used a 120 s timeout. The Semantic category, in which TiPToP's margin is largest (26/40 vs. 10/40), consists entirely of unmarked scenes; AirPods→cup is also unmarked. Since π0.5 is a closed-loop policy that may need multiple grasp attempts and recovery cycles, an 800-step cutoff can convert a late success into a failure precisely on these semantic and multi-step tasks. Appendix C asserts the limits are 'generous' but gives no π0.5 completion-time distribution, and Table II reports only mean time-to-success on successful trials, which cannot establish that the cutoff was non-binding. Please rerun the external scenes with a matched wall-clock budget (120 s), or provide time-to-success and time-to-failure di","section":"Appendix C; Table I"},{"comment":"Real-world scenes use 5 trials per scene and no error bars, confidence intervals, or significance tests. Many per-scene differences, such as 1/5 vs. 4/5 or 2/5 vs. 5/5, are within binomial noise. Task selection was explicitly based on 'tasks that both TiPToP and π0.5-DROID seemed capable of,' and the tasks were then grouped into categories that reward TiPToP's symbolic grounding and long-horizon planning strengths. This is not necessarily invalid, but the paper should justify the menu and provide per-category confidence intervals or a permutation test over scenes. Without this, the aggregate 74.6% vs. 52.4% cannot be cleanly separated from task-selection effects.","section":"Section VII-A; Table I"},{"comment":"The abstract claims that TiPToP 'ranks first overall on pick and pick-and-place tasks among methods not trained on in-distribution data' on the MolmoSpaces benchmark. No section of the paper describes MolmoSpaces, the comparison set, the task suite, or the numeric results. A benchmark-ranking claim cannot be evaluated from the abstract alone. Either add the full MolmoSpaces evaluation to the experiments or remove the claim from the abstract.","section":"Abstract"},{"comment":"The completion-time comparison (Q2) uses average time-to-success over successful trials only, with a manually stopped timer for π0.5 and an automatically stopped timer for TiPToP, and no per-trial distributions or sample sizes are reported. Because failed π0.5 trials can be long and are excluded, the five-of-six speed advantage may overstate the difference. Please report all trials or medians/ranges, and use the same termination and measurement procedure for both systems.","section":"Table II; Appendix C"}],"minor_comments":[{"comment":"There are typographical errors: 'Open-V ocabulary' in the title and 'logical relations betweeen' in Section IV-B. Please proofread the manuscript.","section":"Title; Section IV-B"},{"comment":"The impedance controller equation includes gains Kp and Kd, and the text says they were tuned, but no numerical values or tuning procedure are given. Provide the values in the appendix or point to the open-source controller for reproducibility.","section":"Appendix B"},{"comment":"Task-progress metrics are defined per-task with different scoring rules and penalties, and Appendix C notes that progress metrics 'may vary by the evaluator and the task.' Aggregating these heterogeneous scores in the TP column of Table I is not meaningful across scenes. Report TP only within matched scenes or provide a consistent metric.","section":"Appendix C; Table I"},{"comment":"The table reports mean time-to-success without indicating the number of successful trials used for each mean. Some entries are likely based on a single success. Include per-trial values or at least counts and confidence intervals.","section":"Table II"},{"comment":"The abstract says TiPToP can be deployed on a standard DROID setup in under an hour, while the UR5e adaptation is described as taking 'a few hours.' Clarify that the sub-hour figure applies only to the already-supported DROID configuration, not to new embodiments.","section":"Section VII-C"}],"recommendation":"major_revision","confidential_remarks":"The paper is a credible candidate for publication if the evaluation-protocol issues are fixed. The timeout asymmetry for π0.5-DROID is the most serious technical concern; the MolmoSpaces claim must also be substantiated. I do not see a fundamental correctness error in the system design itself, and the open-source release plus external evaluation are strong selling points."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a legitimate systems contribution. The components are all published, but the integration is new and the evaluation is substantial: 165 trials, 28 scenes, an external deployment team, and a head-to-head against pi0.5-DROID. The aggregate margin (74.6% vs 52.4%) is large and consistent across distractor and multi-step categories, not just one favorable subset. The component-level failure analysis is genuinely useful, and shipping the code makes this a reproducible baseline the community can actually build on.\n\nThe soft spots are real but mostly fixable. The task menu was chosen by the authors based on what both systems \"seemed capable of,\" real-world scenes have only 5 trials with no error bars, and termination limits differ between external (800 steps) and internal (120 s) runs. The stress-test's concern about the step cap is worth taking seriously: a reactive policy that gets cut off at ~53 seconds could underperform on exactly the semantic tasks where TiPToP's margin is largest. But it doesn't sink the comparison. Restricting to external-only scenes, TiPToP still leads 42-27 over 75 trials, so a matched-timeout rerun is needed but unlikely to flip the overall picture. A bigger annoyance is the abstract's MolmoSpaces \"ranks first\" claim, which appears nowhere in the body; that should be supplied or removed.\n\nThe paper is also honest about its main limitation: open-loop execution, explicitly acknowledged and responsible for most failures (grasping). The failure analysis is run on separate tasks from the benchmark, which is fine for debugging but not a substitute for benchmark evidence.\n\nSerious thinker: yes. The paper is coherent on its own terms, engages the relevant literature, and doesn't overclaim beyond its actual evidence—except for the MolmoSpaces sentence. It deserves a serious referee, not a desk reject. Send it to review, but the revision should require matched timeouts, confidence intervals or raw trial logs, and either MolmoSpaces results or a retraction of the claim. I would bring it to a reading group and cite it as a baseline.","headline":"A credible zero-training modular baseline that beats a fine-tuned VLA on a self-selected benchmark; the evaluation protocol needs tightening before the headline is taken at face value, but the core result is probably right.","tokens_in":21797,"tokens_out":2653,"would_cite":true,"duration_ms":24848,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A modular planning system with zero robot training data matches or beats a model fine-tuned on 350 hours of demonstrations.","keywords":["robot manipulation","task and motion planning","vision-language-action models","foundation models","open-vocabulary","zero-shot","modular systems","pick-and-place"],"falsifier":"Run both systems on 100 third-party-designed tasks with identical success criteria and per-task time budgets matched to each controller's rates, and check whether the aggregate gap persists; a targeted version is to put a single concave object (a banana) in front of TiPToP and observe whether the convex-hull mesh repeatedly causes grasp or placement failures.","tokens_in":20827,"feed_emoji":"🤖","tokens_out":5431,"duration_ms":46743,"temperature":0.7,"pith_summary":"This paper tries to establish that a capable robot manipulation system can be built entirely from off-the-shelf components, with no robot-specific training data. TiPToP takes a stereo image pair and a natural-language instruction, builds a 3D object-centric scene model using pretrained vision models, grounds the instruction into symbolic predicates, and then lets a GPU-parallelized task-and-motion planner compute a full trajectory. Across 28 scenes and 165 trials, it attains a 74.6% success rate and faster average completion times on five of six measured scenes, compared with 52.4% for a state-of-the-art vision-language-action model fine-tuned on 350 hours of demonstrations. If this holds, it would mean that open-vocabulary, multi-step manipulation is not the exclusive province of end-to-end learned policies, and that modular systems can be debugged, extended, and improved component by component.","feed_headline":"Modular planner beats 350-hour-trained robot model","feed_subtitle":"Zero-training system built from off-the-shelf models matches or tops a fine-tuned VLA across 28 scenes.","key_machinery":"The central mechanism is a two-branch pipeline that fuses semantic and geometric understanding into an object-centric scene representation, then hands it to cuTAMP, a GPU-parallelized task-and-motion planner. A vision-language model turns the instruction into symbolic predicates (currently on(a,b)), grounding open-vocabulary references like 'peanut butter crackers' or 'largest toy' onto detected objects; stereo depth, segmentation, and grasp-prediction models supply per-object meshes and candidate grasps. cuTAMP enumerates plan skeletons, initializes thousands of sampled solutions, and jointly optimizes grasp and placement poses against collision, stability, and kinematic constraints, genera","core_discovery":"The authors claim that composing pretrained depth, segmentation, grasp, and language models with a GPU-parallelized task-and-motion planner (cuTAMP) yields a manipulation system that, with zero robot training data, matches or outperforms π0.5-DROID, a VLA fine-tuned on 350 hours of demonstrations. Across 28 scenes and 165 trials, TiPToP attains 74.6% success versus 52.4%, higher task progress in distractor, semantic, and multi-step categories, and faster time-to-success on five of six scenes. They also show that the modular architecture enables component-level failure tracing, with grasping as the dominant bottleneck, and that new embodiments and skills can be added within hours.","pith_inferences":["The headline comparison is sensitive to the evaluation protocol: tasks were chosen to suit both systems, categories reward VLM grounding and multi-step geometric reasoning, and baseline termination limits differ; an independently curated task set could shrink or reverse the gap.","The open-loop architecture means robustness depends on static scenes and precise tracking; adding closed-loop replanning after each pick-and-place, as the paper itself suggests, would likely address the dominant grasp-failure mode.","The modular decomposition yields a direct testable extension: swapping in a stronger vision-language model should improve semantic and distractor tasks without touching the planner, while swapping in a better grasp predictor should directly reduce the largest failure class.","Single-viewpoint convex-hull meshes are the root of failures on concave objects like bananas; multi-view perception or learned shape completion is a natural next experiment that would test whether the perception module, not the planner, is the binding constraint."],"forward_implications":["A manipulation system that requires no robot training data and can be installed on a standard DROID setup in under an hour is a viable alternative to end-to-end VLAs for pick-and-place and multi-step tasks.","Component-level failure tracing becomes practical, steering improvement effort to the weakest modules—grasping first, then scene completion, VLM detection, and planning.","Time-to-success is roughly half that of the reactive VLA on single-step real-world tasks, because the planner commits to a single time-optimal trajectory instead of iterating a closed-loop policy.","The complementary failure modes suggest a hybrid design: using a VLA as a closed-loop skill primitive inside the TAMP framework to recover from grasp slips and unexpected object motion.","Because components are swappable, the system should improve automatically as better depth estimators, grasp predictors, and VLMs become available."],"fun_headline_variants":["Zero-training robot system beats 350-hour trained model","Modular planner outperforms VLA trained on 350h of data","TiPToP: no robot training data tops fine-tuned VLA","Open-source modular robot system surpasses trained VLA","TiPToP: zero training beats 350-hour VLA"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central premise that carries the headline number is that the chosen evaluation scenes, task menus, and termination limits treat the trained baseline fairly; if those choices instead favor a system that reasons about semantics and geometry, the 74.6% vs 52.4% gap is not a general statement.","fun_headline_variants_meta":{"raw":{"variants":["Zero-training robot system beats 350-hour trained model","Modular planner outperforms VLA trained on 350h of data","TiPToP: no robot training data tops fine-tuned VLA","Open-source modular robot system surpasses trained VLA","TiPToP: zero training beats 350-hour VLA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001411,"raw_usage":{"total_tokens":5551,"prompt_tokens":771,"completion_tokens":4780,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":515,"completion_tokens_details":{"reasoning_tokens":4692}},"tokens_in":515,"tokens_out":4780,"duration_ms":30944,"temperature":1.0,"reasoning_tokens":4692,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T18:26:18.376264+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run both systems on 100 third-party-designed tasks with identical success criteria and per-task time budgets matched to each controller's rates, and check whether the aggregate gap persists; a targeted version is to put a single concave object (a banana) in front of TiPToP and observe whether the convex-hull mesh repeatedly causes grasp or placement failures.","supporting_citations":[],"review_version":1}