{"id":"d302d779-9067-4f4e-8aeb-7936fab7be8d","arxiv_id":"2412.14989","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"NimbRo's TIAGo++ robot won RoboCup@Home 2024 OPL using open-vocabulary segmentation and LLM-based task planning.","lead":"NimbRo's TIAGo++ service robot won the RoboCup@Home 2024 Open Platform League by combining open-vocabulary image segmentation with large language models for planning. The result shows how off-the-shelf AI foundation models can work together on a physical robot for household tasks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Open-vocabulary generalization claim is confounded by on-site prompt and data tuning; the competition win does not isolate the foundation-model contribution.","rationale":"The reader's weakest assumption — that manual prompt design and on-site curated data were the decisive factor — is precisely the load-bearing concern. The paper's own Section 3.3 admits both the manual, locally-dataset-driven prompt design and the collection of additional data for poorly performing objects. This directly undermines the abstract's 'robustness and generalization capabilities' claim, because the system had access to the evaluation environment before the test. The competition win itself is independently supported by the reported scores and is not in question; what is in question is the attribution of that win to open-vocabulary foundation models. No numbers are given for open-vocabulary segmentation accuracy, grasp success, or LLM planning success, and the final demonstration is a single, unrepeated showcase. The paper is honest about these limitations, which is commendable, but honesty about the confound does not remove it. Given that the reader already issued a conditional verdict identifying this exact issue, my stress-test does not change the verdict; it reinforces the need for the paper to either release artifacts and ablations or restrict its claims to the demonstrated competition outcome. The proposed concrete test — a frozen-prompt, no-additional-data evaluation in a new arena — would directly settle whether the open-vocabulary generalization claim is empirically supported or whether the win reflects prompt/data tuning plus robust system integration.","tokens_in":8267,"tokens_out":3320,"duration_ms":32870,"concrete_test":"Release the exact prompt set, the on-site dataset additions, and the configuration logs, then run a post-competition evaluation in a new, unseen arena with prompts frozen before arena exposure and no additional data collection. Compare mmGrounding-DINO + NanoSAM grasp success and task scores against the supervised MaskDINO pipeline under otherwise identical conditions. If frozen-prompt open-vocabulary performance is comparable, the generalization claim stands; if it drops substantially, the decisive factors were on-site prompt/data tuning and closed-set components, and the paper should be read as an engineering result rather than evidence of open-vocabulary generalization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim — that open-vocabulary segmentation and LLM planning enabled segmenting and grasping non-labeled objects with robustness and generalization — is underdetermined by the evidence presented. Section 3.3 states the open-vocabulary prompts 'were manually designed using our locally captured dataset for evaluation and could be changed on the fly during the task,' and that 'over the setup days, more data was collected for objects which were not performing well.' This means the 'open-vocabulary' system was tuned to the actual evaluation arena and objects before and during the test. The same section also notes that open-vocabulary models were used only when there were 'few relevant object classes (Serving Breakfast, Clean the Table),' while other tasks relied on fine-tuned closed-vocabulary MaskDINO/YOLO models with per-task data collection. The overall win, therefore, could be driven by the supervised components, SLAM, person tracking, grasp planning, and system engineering rather than by foundation-model generalization. The paper reports only aggregate task scores and a single final-demonstration anecdote; there are no per-component metrics such as segmentation IoU, grasp success rate, or LLM planning success. Thus the paper's own limitations text, combined with the absence of ablations, supports the reader's concern: the headline claim about open-vocabulary generalization is not falsifiable from the reported data. This is not an internal inconsistency, but a causal-attribution gap between the verified competition outcome and the broader scientific claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports the approaches, hardware, and results of team NimbRo@Home at the RoboCup@Home 2024 Open Platform League, where the team won first place. The main technical emphasis is on open-vocabulary object segmentation and grasping, and on the use of large language models (LLMs) for natural language understanding and task planning. The paper describes the perception, planning, grasping, speech, and navigation modules, reports per-task competition outcomes and scores for the three stages and the final demonstration, and concludes that open-vocabulary segmentation proved valuable and that robustness and generalization were key to the win.","tokens_in":8639,"tokens_out":4683,"duration_ms":30011,"significance":"If the central claim is accepted, the paper demonstrates a practical integration of foundation models (open-vocabulary segmentation and LLM planning) into a deployed service robot that won a major competition. This would be a notable existence proof that promptable foundation models can reduce task-specific supervision in domestic robotics. The competition outcome itself is externally documented and the paper provides a useful system-level integration description, including a video reference and a comparison of final scores with other teams. However, the strength of the scientific contribution is limited by the absence of per-component quantitative evaluation and by the acknowledged on-site tuning of prompts and datasets, which weakens the generalization and robustness claims as stated.","major_comments":[{"comment":"The claim in the abstract that open-vocabulary segmentation and grasping of 'non-labeled objects' was demonstrated with 'robustness and generalization capabilities' is substantially weakened by the description of on-site tuning. The text states that prompts 'were manually designed using our locally captured dataset for evaluation and could be changed on the fly during the task' and that 'over the setup days, more data was collected for objects which were not performing well.' This means the open-vocabulary pipeline was adapted to the evaluation arena and its specific objects before and during the tests. The paper should either report which prompts were fixed a priori and which were changed, provide a held-out evaluation on objects not seen during the tuning days, or explicitly temper the generalization claim in the abstract and conclusion.","section":"Section 3.3"},{"comment":"The quantitative evidence consists solely of aggregate competition scores and single-run task anecdotes. No per-component metrics are reported, such as segmentation intersection-over-union, grasp success rates, number of grasp reattempts, LLM planning success rates, or a comparison between the open-vocabulary and closed-vocabulary pipelines on the same tasks. Consequently, the paper does not support the attribution of the competition win to the open-vocabulary and LLM components rather than to other subsystems (e.g., SLAM, person tracking, touchscreen fallback, closed-vocabulary detectors, or the overall system engineering). A table reporting per-task component success/failure counts and, where available, the scores from both arena runs would make the contribution of each component assessable.","section":"Section 4"},{"comment":"The final demonstration (egg pouring) is a single anecdote used to support the headline claim that open-vocabulary approaches can grasp non-labeled objects and execute complex tasks. Single demonstrations, without repeated trials, failure counts, or any quantitative measure of perception or manipulation success, do not support statements of robustness or generalization. The paper should report how many objects were scanned, how many grasps were attempted and succeeded, and how the open-vocabulary perception output was consumed by the task planner, or it should restrict the claim to a feasibility demonstration.","section":"Section 4.3"},{"comment":"The Lessons Learned section states that 'Open-vocabulary instance segmentation proved valuable in this competition,' but open-vocabulary models were used in only two stage tasks (Serving Breakfast and Clean the Table) and in combination with closed-vocabulary models in other tasks (e.g., Stickler for the Rules). Without a direct comparison or per-task attribution of the open-vocabulary contribution, this conclusion is not supported by the reported data. Please either provide such evidence or qualify the lesson to reflect the actual scope of open-vocabulary usage.","section":"Section 5"}],"minor_comments":[{"comment":"The degree symbol in '180◦ FOV' should be typeset as '180°'; check for similar typographical issues throughout the paper.","section":"Section 2"},{"comment":"The phrase 'MeanIntersectionoverUnion' should be expanded with proper spacing as 'Mean Intersection over Union' for readability.","section":"Section 3.3"},{"comment":"The abbreviation 'SOTA' should be expanded on first use (e.g., 'state-of-the-art models') to make the text self-contained.","section":"Section 3.6"},{"comment":"The paper notes that 'the tests were executed twice in different arenas' but reports only a single aggregate score per stage. Reporting the per-run scores would give readers a better sense of variability.","section":"Section 4"},{"comment":"The object perception pipeline diagram is visually dense; consider enlarging the figure or separating it into two panels for legibility.","section":"Figure 4"},{"comment":"Several reference entries for online resources (e.g., JACK Audio, Coqui TTS, Faster Whisper) lack access dates; add consistent access-date information where applicable.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a competition team-description paper, and the external ranking provides solid evidence for the win. The main gap is the mismatch between the broad claims in the abstract and the limited, heavily tuned evidence presented. The paper would be acceptable after a major revision that either adds per-component quantitative evaluation (even simple success/failure counts) or substantially softens the generalization and robustness claims. The absence of ablations is the key scientific weakness; the authors should be encouraged to make the on-site tuning explicit in the main text rather than only in the pipeline description."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nTwo things to know. First, this is a genuine systems win: NimbRo@Home finished first in RoboCup@Home 2024 OPL with a clear margin over Tidyboy-OPL, and the scores are externally documented. Second, the headline scientific claim — that open-vocabulary segmentation and LLM planning generalize to unseen household objects — is not isolated by the evidence. The stress-test note is correct. Section 3.3 states the open-vocabulary prompts were manually designed on a locally captured dataset and 'could be changed on the fly during the task,' and more data was collected for objects that performed poorly. So the system was tuned to the actual arena before and during the test. The same section says open-vocabulary models were used only in tasks with few relevant classes (Serving Breakfast, Clean the Table); other tasks used fine-tuned MaskDINO/YOLO models with per-task data collection. The overall win could easily be driven by the supervised components, SLAM, tracking, grasp planning, and system integration. There are no per-component metrics — no segmentation IoU, grasp success rate, or LLM planning success — and no ablations. That is a real attribution gap between the verified competition result and the robustness/generalization language in the abstract.\n\nWhat the paper does well: it is an honest, well-scoped team description. Hardware details are concrete, the software modules are clearly described, and results match the published standings. The hybrid approach — combining closed-set detectors with open-vocabulary models for categories like drinks, shoes, garbage — is a sensible engineering choice, described candidly. The LLM function-calling framework for GPSR/EGPSR, including rejection of infeasible commands, is a useful implementation detail. The citation pattern is fine: prior NimbRo work is context, and the competition result itself is the evidence; no fitted parameters are used to derive the outcome, so circularity is not an issue.\n\nWho gets value: readers who want to see how a top RoboCup@Home team integrates foundation models with classical robotics will find this useful. Readers looking for falsifiable evidence about open-vocabulary generalization will not find it here. It deserves a serious referee — the competition win and integration work are worth peer review — but the referee should require a clearer separation of the verified outcome from the unverified generality claims, and ideally release prompts, configurations, and run logs. I would not cite this as evidence of foundation-model robustness, but I might cite it as an example of successful system integration in service robotics. Recommend: send to peer review with a requested revision that separates outcome from generalization claims.","headline":"Real competition win, externally scored; but the open-vocabulary generalization claim is not isolated from on-site tuning — a solid systems paper that needs clearer separation of verified outcome from broader claims.","tokens_in":9094,"tokens_out":3176,"would_cite":false,"duration_ms":19500,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper reports that a service robot combining open-vocabulary object segmentation with LLM-based task planning won the RoboCup@Home 2024 Open Platform League, showing that foundation models can replace much task-specific supervision in…","keywords":["service robotics","open-vocabulary object segmentation","grounded language-image pretraining","LLM task planning","mobile manipulation","RoboCup@Home","foundation models","promptable segmentation"],"falsifier":"Run the same robot in a fresh apartment using only generic object names with no on-site dataset and no prompt tuning; if it cannot segment and grasp a held-out set of household objects such as a mug, a sponge, and a pear with comparable success, the paper's claim that open-vocabulary approaches overcame labeling overhead is not supported.","tokens_in":8082,"feed_emoji":"🤖","tokens_out":9790,"duration_ms":58343,"temperature":0.7,"pith_summary":"At the RoboCup@Home 2024 Open Platform League, the authors' team fielded a service robot whose perception combined a fine-tuned closed-vocabulary detector with an open-vocabulary grounding model that segments objects from free-text descriptions. The paper's central claim is that this hybrid foundation-model pipeline let the robot segment and grasp objects it had never been labeled for, and that an LLM using function calling turned natural-language commands into multi-step tasks. The team won the overall competition, taking top scores in the GPSR, Restaurant, and final-demonstration stages. If the claim is right, it shows that text-promptable perception plus LLM planning can reduce the labeling overhead that currently dominates domestic service robotics.","feed_headline":"Robot wins RoboCup@Home 2024 by grasping unseen objects","feed_subtitle":"Text-promptable vision and an LLM planner handled household objects the robot had not been trained on.","key_machinery":"The load-bearing pairing is text-promptable grounding plus promptable segmentation: mmGrounding-DINO, an open-vocabulary detector that finds objects from text descriptions, outputs bounding boxes, and NanoSAM, a lightweight promptable segmenter, converts those boxes into instance masks. Masks are projected into depth to build a partial point cloud, optionally completed by registering a 3D model, and approximated by an oriented bounding box. Grasp poses are sampled on a quadrant sphere facing the object, filtered for collisions using a KD-tree over composite RGB-D and LiDAR point clouds, and ranked by a heuristic that favors precomputed collision-free pre-grasp poses with clearance from obstacles and the workspace boundary. Task planning is carried by an LLM that calls robot capability functions with textual feedback, allowing about 3 to 15 function calls per typical command.","core_discovery":"The authors report that open-vocabulary object segmentation is practically usable in a competitive household-robot setting: a text-promptable grounding detector (mmGrounding-DINO) produced bounding boxes for objects named in natural language, and a promptable segmenter (NanoSAM) refined those boxes into instance masks that fed the grasping pipeline. The same perception channel provided semantic scene information in the final demonstration, where the robot scanned an apartment, identified present objects, and later poured an egg into a pan. An LLM (GPT-4o) executed natural-language commands by calling a library of robot capability functions, advancing a task state machine until the command was accomplished or rejecting commands outside the robot's abilities. The authors state that this approach, together with a supervised detector for known object classes, won the Open Platform League with a total score of 8,852 points.","pith_inferences":["A natural extension the paper does not itself claim is that this architecture transfers to other domestic chores, such as tidying arbitrary objects or fetching items described by appearance, because the perception and planning layers are not tied to the competition's object classes.","The on-site prompt and dataset tuning described in Section 3.3 suggests a testable boundary: if only generic prompts and no locally captured data are allowed, performance on unseen objects may drop, which would show that some of the claimed generalization is task-specific adaptation.","The GPSR command about the smallest object on a counter hints that the same LLM-plus-vision loop could answer property questions (color, material, size) without a bespoke attribute classifier, by grounding attributes through text prompts and visual segments.","A direct end-to-end measurement with prompts frozen at the start of a competition, rather than adjustable on the fly, would separate the contribution of the open-vocabulary models from the contribution of human prompt engineering."],"forward_implications":["Household robots could be deployed to new environments without collecting and labeling a task-specific object dataset for every home, since text descriptions can stand in for training examples.","LLM function calling with textual feedback can handle varied natural-language commands, including rejecting requests the robot cannot or should not perform, which matters for non-expert users.","Combining closed-set detectors for known categories with open-vocabulary models for unknown items (shoes, socks, drinks, garbage) improves both precision and recall in monitoring tasks like Stickler for the Rules.","The same perception pipeline carried over to the final demonstration, scanning a kitchen, supporting user-input-based planning, and executing a complex manipulation such as pouring an egg into a pan, suggesting reuse beyond predefined task stages."],"supporting_citations":[{"why":"Supplies the open-vocabulary grounding detector mmGrounding-DINO that produces bounding boxes from text prompts.","marker":"[33]"},{"why":"Segment Anything, used for semi-automatic annotation of the training dataset and as the basis for NanoSAM's mask generation.","marker":"[11]"},{"why":"GPT-4o, the LLM accessed through the robot's hybrid Wi-Fi/5G link for natural-language understanding and function-calling task planning.","marker":"[19]"},{"why":"cuRobo, used to precompute the reachability map that anchors grasp-pose sampling for manipulation.","marker":"[30]"},{"why":"Mask DINO, the fine-tuned instance segmentation model used as the closed-vocabulary alternative in the hybrid perception pipeline.","marker":"[14]"},{"why":"YOLOv8, used for person pose and face detection and for closed-set object detection in the same pipeline.","marker":"[9]"},{"why":"A comparison of prompt engineering techniques that informed how the LLM is prompted for task planning and execution.","marker":"[3]"},{"why":"The RoboCup@Home 2024 rule book, which defines the task stages and scoring that the reported results are measured against.","marker":"[8]"}],"fun_headline_variants":["NimbRo wins RoboCup@Home 2024 with open-vocab grasping","LLM and open-vocab vision win RoboCup@Home 2024","Language prompts let robot grab unseen objects, win RoboCup@Home","NimbRo's open-vocab vision and LLM win RoboCup@Home","Open-vocab grasping wins RoboCup@Home 2024 for NimbRo"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The open-vocabulary generalization claim rests on the assumption that the manually designed text prompts and the on-site dataset used for tuning were not the decisive factor, since Section 3.3 says prompts were designed using locally captured data and could be changed on the fly, and additional data was collected for objects that performed poorly.","fun_headline_variants_meta":{"raw":{"variants":["NimbRo wins RoboCup@Home 2024 with open-vocab grasping","LLM and open-vocab vision win RoboCup@Home 2024","Language prompts let robot grab unseen objects, win RoboCup@Home","NimbRo's open-vocab vision and LLM win RoboCup@Home","Open-vocab grasping wins RoboCup@Home 2024 for NimbRo"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00102,"raw_usage":{"total_tokens":4260,"prompt_tokens":855,"completion_tokens":3405,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":471,"completion_tokens_details":{"reasoning_tokens":3294}},"tokens_in":471,"tokens_out":3405,"duration_ms":19326,"temperature":1.0,"reasoning_tokens":3294,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:43:20.167547+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same robot in a fresh apartment using only generic object names with no on-site dataset and no prompt tuning; if it cannot segment and grasp a held-out set of household objects such as a mug, a sponge, and a pear with comparable success, the paper's claim that open-vocabulary approaches overcame labeling overhead is not supported.","supporting_citations":[{"cited_title":"In: IEEE/CVF International Conference on Computer Vision (ICCV)","cited_arxiv_id":null,"evidence_quote":"Segment Anything, used for semi-automatic annotation of the training dataset and as the basis for NanoSAM's mask generation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"GPT-4o, the LLM accessed through the robot's hybrid Wi-Fi/5G link for natural-language understanding and function-calling task planning."},{"cited_title":"In: IEEE International Conference on Robotics and Automation (ICRA)","cited_arxiv_id":null,"evidence_quote":"cuRobo, used to precompute the reachability map that anchors grasp-pose sampling for manipulation."},{"cited_title":"In: IEEE/CVF Conf","cited_arxiv_id":null,"evidence_quote":"Mask DINO, the fine-tuned instance segmentation model used as the closed-vocabulary alternative in the hybrid perception pipeline."},{"cited_title":"https://github.com/ ultralytics/ultralytics (2023)","cited_arxiv_id":null,"evidence_quote":"YOLOv8, used for person pose and face detection and for closed-set object detection in the same pipeline."},{"cited_title":"In: IEEE- RAS 23rd International Conference on Humanoid Robots (Humanoids) (2024)","cited_arxiv_id":null,"evidence_quote":"A comparison of prompt engineering techniques that informed how the LLM is prompted for task planning and execution."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The RoboCup@Home 2024 rule book, which defines the task stages and scoring that the reported results are measured against."}],"review_version":1}