{"id":"8222a541-ebed-49d9-ba8e-dbb838faace9","arxiv_id":"2412.20998","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"T-DOM is a taxonomy for deformable object manipulation that adds a force-direction-based deformation classification, including new structured and unstructured bending levels, evaluated on ten curated tasks.","lead":"The paper introduces T-DOM, a taxonomy that sorts robotic manipulation of deformable objects by deformation type, motion energy, and contact interactions. It also proposes classifying bending as structured or unstructured, and tests the taxonomy on ten recorded manipulation tasks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim rests on subjective manual deformation labels; without inter-annotator agreement or independent physical verification, the cluster-based differentiation in §5.3 may reflect annotation choices rather than objective task structure.","rationale":"The reader's weakest_assumption exactly identifies the most load-bearing concern: the manual labels in §5.2 are the sole input to the cluster comparison, yet they are never tested for reliability or physical grounding. I agree with the reader's CONDITIONAL verdict because the taxonomy itself is plausible and clearly presented, and the proposed deformation categories are genuinely useful as an organizing vocabulary even if the current evidence is thin. The concern does not move the verdict because the reader already conditioned acceptance on stronger validation; it does, however, sharpen why that validation is essential. I see no reason to upgrade to ACCEPT without the check, and no reason to reject outright: the paper's contribution is a taxonomy whose value is not fully falsified by weak evaluation, and the authors openly position the dataset as open-source and the analysis as descriptive. The weakest part of the argument is precisely the assumption that human-annotated deformation labels are objective enough to support the claim that deformation classification is essential. Since the clustering procedure is definitional, the only way the comparison can be evidence is if the tags themselves are reliable. Independent annotation agreement alone would not prove task-relevant information, but it is the minimal check needed before interpreting Fig. 5 as evidence rather than as a restatement of the taxonomy's own definitions.","tokens_in":25376,"tokens_out":5511,"duration_ms":61410,"concrete_test":"Recruit two or more independent annotators who have not seen Table 2; give them the taxonomy definitions from §3–§4 and the same RGB-D task videos; ask them to segment each of the 10 tasks and assign full T-DOM tags. Compute Cohen's kappa (or Fleiss' kappa) per category, especially Deformation, S-level, and US-level. Additionally, for at least the tension/compression rows (e.g., T3-4, T5-3), measure physical deformation using force sensing or optical strain tracking and check whether the assigned deformation category matches the measured strain direction and magnitude. If kappa is below 0.6 or measured strain contradicts the labels, the differentiation in Fig. 5 likely reflects annotation choices rather than task structure, and the central claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The empirical support for the central claim in §5.3 is built almost entirely from clustering action-IDs constructed from the authors' own manual labels (Table 2; §5.2). Because clusters are defined by equality of tag vectors, removing any category, including deformation, is mathematically guaranteed to coarsen or merge clusters; the ablation in Fig. 5b therefore only shows that the deformation column is not redundant, not that it is semantically meaningful. Meaningfulness requires that the labels be reproducible and correspond to measurable physical deformation. The paper reports no inter-annotator reliability, no annotation protocol, and no independent verification. For categories such as compression in a folded towel, tension versus bending in edge tracing, and structured versus unstructured bending levels, the boundary decisions are non-obvious: S L0/L1/L2 and US L0/L1/L2 are qualitative counts of loops, g-folds, or visible keypoints, and Table 2 even contains ambiguous rows (e.g., rows 5-4 and 5-5 shift from S L0 to N without explanation). If a second annotator would assign different deformation tags, the claimed differentiation across similar robot actions is not robust. This is not a fatal flaw — a taxonomy can be useful despite subjective labels — but the paper's cluster-based evaluation cannot carry the central claim without validation of the labels themselves.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces T-DOM, a taxonomy designed to describe the robotic manipulation of deformable objects. T-DOM has three main branches: deformation (compression, tension, torsion, shear, and structured/unstructured bending), robot motion (kinetic, elastic, gravitational, and gravitational-elastic energy dominance), and interactions (prehensile point/line grasps, non-prehensile agent/environment contacts, and contact sliding). To evaluate the taxonomy, the authors record ten manipulation tasks, segment them into individual actions, label each action with tag vectors from T-DOM, and compare the resulting clustering of action-IDs against those induced by the taxonomies of Bullock et al. (2012) and Paulius et al. (2020). An ablation that removes the deformation column from the tag vectors is used to argue that deformation is essential for distinguishing similar robot actions.","tokens_in":25608,"tokens_out":5388,"duration_ms":55697,"significance":"If the taxonomy and its evaluation are robust, T-DOM would fill a genuine gap: existing manipulation taxonomies are mostly designed around rigid objects, and the deformable-manipulation community lacks a shared descriptive vocabulary. The proposed structured/unstructured bending distinction, the energy-based motion categories, and the decoupled bimanual tagging scheme are sensible and potentially useful design choices. The ten-task dataset, with open RGB-D data and explicit action-ID tables, is a concrete and reusable contribution. The comparison with Bullock et al. and Paulius et al. is fair in acknowledging that those taxonomies had different goals. However, the empirical support for the central claim is currently weak. The ablation in Fig. 5b is logically guaranteed to coarsen the clustering when any column is removed, so it does not by itself establish that deformation is semantically meaningful. The manual labels in Table 2 have no reported inter-annotator reliability, annotation protocol, or quantitative grounding. These issues are load-bearing because the paper's main conclusion rests on the action-ID clusters and the ablation.","major_comments":[{"comment":"The ablation that removes the deformation category does not support the claim that deformation is essential. Since an action-ID is an ordered vector of tag values, deleting any one column can only merge previously distinct IDs; it cannot create new distinctions. The coarsening seen in Fig. 5b is therefore a mathematical consequence of the representation, and it only shows that the deformation column is not redundant in this particular dataset. To establish semantic meaningfulness, the paper would need to show that the deformation labels are reproducible (e.g., through inter-annotator agreement) or correlate with independently measurable physical quantities such as strain, curvature, or topological descriptors of the kind mentioned in §6.2. Without such evidence, the cluster separation in Fig. 5a may reflect the annotators' choices rather than objective deformation states.","section":"§5.3, Fig. 5b"},{"comment":"The evaluation rests entirely on manual segmentation and manual tag assignment, with no inter-annotator reliability, no annotation protocol, and no independent verification. Several boundary decisions are non-obvious: the S L0/L1/L2 and US L0/L1/L2 levels are defined by qualitative counts of loops, g-folds, or visible keypoints, and Table 2 contains unexplained transitions such as rows 5-4 (S L0) to 5-5 (S N) in the Transport meat task and the simultaneous changes in both S and US between rows 6-3 and 6-4 of the Flatten cloth task. The paper should provide concrete annotation guidelines, ideally a second annotator's independent labels, and a description of how disagreements were resolved. This is essential because the central claim in §5.3—that classifying deformation is essential for distinguishing similar robot actions—is only as strong as the reliability of these labels.","section":"§5.2, Table 2"},{"comment":"The qualitative bending scale is not sufficiently specified to be applied by others. For 1D objects, the levels are tied to loops and removable crossings; for 2D objects, to g-folds and accessible keypoints; but for 3D objects, structured bending is defined only as 'measured in the plane curve where the bending deformation takes place', with no operational procedure for identifying that curve or counting levels. The boundary between S L0, S L1, and S L2, and between US L0, US L1, and US L2, is described through examples rather than rules. Because the deformation labels are the main differentiator in the evaluation, the taxonomy needs precise, reproducible definitions of these levels before the experimental claims can be fully assessed.","section":"§3.1, Fig. 2"}],"minor_comments":[{"comment":"There is a typo: 'slidinf' should be 'sliding', and the sentence 'Some examples include and end-effector grasping and slidinf along an edge of a cloth' should read 'an end-effector grasping and sliding'. Additionally, the definition of active sliding says it occurs 'without grasping it', but the example describes an end-effector grasping and sliding along an edge; this inconsistency should be clarified.","section":"§4.5"},{"comment":"The phrase 'for the first time, a detailed classification of object deformations' is stronger than necessary, since Paulius et al. (2020) already includes a temporal/permanent deformation distinction and rigid/soft engagement. The novelty is better stated as the classification of deformation type by force direction, not the first deformation classification of any kind.","section":"Abstract and Introduction"},{"comment":"The diagram uses AND/OR logic and a dashed line for contact sliding, but the caption does not explain how to read these elements. Please add a short legend or caption explanation, since the figure is the key reference for interpreting the tag vectors in Table 2.","section":"Figure 3"},{"comment":"The table header abbreviations (N-P.Env., N-P.Act., CS, S, US) are not expanded in the caption or table notes. Readers must cross-reference Figure 3 and Section 4, which is cumbersome. Please add full names or a legend.","section":"Table 2"},{"comment":"When the paper states that unimanual tasks use a 'none' tag for the second manipulator, it should clarify whether 'none' is treated as a distinct value in the action-ID equivalence relation or as a wildcard. This matters for the clustering comparisons in Fig. 5.","section":"§5.2"},{"comment":"The discussion of deformation sensing mentions quantitative metrics such as the Gauss linking integral but does not connect them to the S/US levels used in T-DOM. A short paragraph on which existing quantitative metrics could operationalize or validate the qualitative bending levels would strengthen the paper's roadmap.","section":"§6.2"}],"recommendation":"major_revision","confidential_remarks":"This paper addresses a relevant gap in the deformable manipulation literature and offers a plausible taxonomy plus a reusable dataset. However, the evaluation as written does not yet support the central empirical claim. The ablation in Fig. 5b is a redundancy check, not a semantic validation, and the manual labels lack reliability and quantitative grounding. I would encourage the editor to send the paper back for a major revision in which the authors add annotation reliability evidence, provide a more careful interpretation of the ablation, and sharpen the definitions of the structured/unstructured bending levels. The core taxonomy idea is worth another round; the current form is not yet convincing on its own terms."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"T-DOM is a genuine attempt to give deformable object manipulation a shared vocabulary, and the force-direction deformation axis plus the structured/unstructured bending distinction are actually new in a manipulation taxonomy. The paper is worth reading for that alone. But the central claim—that deformation labels are essential for distinguishing similar robot actions—is not supported by the evidence as presented, because the entire analysis rests on the authors' own manual labels with no inter-annotator reliability and no quantitative deformation measures.\n\nWhat is good: the taxonomy is coherent. The deformation categories draw on standard mechanics (compression, tension, bending, torsion, shear), and the bending levels are a concrete attempt to capture object state rather than just forces. The motion-energy categories are basically Mason's dynamic/quasi-static split recast in energy terms, which is fine, and the interaction categories cover non-prehensile environment contacts, an omission in Bullock's taxonomy that matters for placing and folding. The dataset and full tag table in Table 2 are a useful resource, and the comparison against Bullock and Paulius is honest about what each taxonomy was designed to do. The citation pattern is reasonable; prior work by Mason, Borras, Paulius, and others is appropriately credited.\n\nSoft spots: the evaluation in §5.3 is the weakest part. Action-IDs are built directly from the authors' manual labels, so the cluster separation is partly built in. Removing the deformation column is mathematically guaranteed to coarsen clusters; the ablation only shows the column is not redundant, not that it is semantically meaningful. Meaningfulness requires reproducible labels. None are shown. Table 2 has genuinely ambiguous rows—e.g., T5-4 and T5-5 shift from S L0 to no bending without explanation, and the S L1/L2 and US L1/L2 boundaries are qualitative counts. No shearing task is included, which is acknowledged but still limits coverage. No code is released, only the labels.\n\nNone of this kills the taxonomy. A taxonomy can be useful without quantitative validation, and the authors frame this as a first step. But the paper currently overclaims when it says the analysis demonstrates the importance of deformation. That needs either inter-annotator agreement, quantitative deformation metrics, or a more careful statement of what the cluster analysis can and cannot show.\n\nThe paper is for researchers in deformable object manipulation who want a common descriptive language for benchmarks, gripper design, and skill transfer. It deserves serious peer review and likely revision. I would send it to a good robotics venue with clear requests: add inter-rater reliability, address the ambiguous labels, and soften or rework the central claim.","headline":"Useful taxonomy with a genuinely new deformation axis, but the central claim rests on unvalidated manual labels—still worth a serious referee.","tokens_in":26246,"tokens_out":2351,"would_cite":true,"duration_ms":23398,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T40"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces T-DOM, a taxonomy that adds deformation type to robot motion and interaction labels, and argues this split is essential for telling manipulation skills apart.","keywords":["deformable object manipulation","manipulation taxonomy","deformation classification","structured bending","unstructured bending","action segmentation","prehensile and non-prehensile interaction","robot motion taxonomy"],"falsifier":"Have two independent annotators label the same ten task videos with T-DOM. If they agree poorly on where action boundaries lie, or if an automated deformation measure (such as the derivative of the Gauss Linking Integral) fails to change exactly when the manual deformation labels change, the claim that deformation categories objectively differentiate the skills would be weakened.","tokens_in":25047,"feed_emoji":"🤖","tokens_out":6665,"duration_ms":59827,"temperature":0.7,"pith_summary":"Robotic manipulation taxonomies have mostly assumed objects are rigid, so they classify grasps and motions but ignore what happens to the object itself. T-DOM adds a deformation dimension—compression, tension, torsion, shear, and a structured versus unstructured split for bending—alongside robot motion and interaction categories. The paper claims that this deformation information is essential for telling apart manipulation skills that otherwise look identical, and supports the claim by labelling ten manipulation tasks and clustering actions by their taxonomy tags. If the taxonomy is right, it gives the field a shared vocabulary for describing deformable-object tasks, with uses in gripper design, skill learning, and benchmarking.","feed_headline":"Adding deformation labels splits robot skills old taxonomies merge","feed_subtitle":"Adding the deformation dimension changes which robot actions count as the same skill.","key_machinery":"The machinery is a three-part tag system: deformation (D), motion (M), and interactions (G for prehensile grasp, NP for non-prehensile, CS for contact sliding). Each manipulation action is written as an ordered set of tags, an action-ID, and actions sharing the same ID are connected in a graph; the clusters of that graph show which skills a taxonomy treats as interchangeable. The load-bearing novelty inside the tag system is the deformation axis, especially the qualitative bending scale that separates ordered folds (counted by loops in 1D objects or g-folds in 2D objects) from disordered wrinkles and knots (counted by removable crossings or accessible keypoints).","core_discovery":"The central claim is that the type of deformation an object undergoes is task-relevant information that prior taxonomies omit, and that encoding it changes which robot actions count as the same. T-DOM classifies deformation by the direction of applied forces—compression, tension, bending, torsion, shear—introduces structured bending (ordered folds, counted by loops or g-folds) versus unstructured bending (wrinkles and knots, counted by removable crossings or accessible keypoints), and allows combined tags such as tension-plus-torsion when wringing a towel. On a dataset of ten tasks with towels, meat phantoms, gowns, gloves, and cables, actions labelled with T-DOM form smaller, semantically cleaner clusters than the same actions labelled with two established taxonomies, and removing the deformation category merges many distinct grasp actions into one cluster. The paper concludes that deformation labels are necessary to distinguish similar robot actions across different deformable objects.","pith_inferences":["The deformation labels could be predicted automatically from RGB-D or tactile data, which would turn the manual segmentation into a learned perception task and test whether the categories are objectively recoverable.","The action-ID representation suggests a natural interface to language models: a T-DOM tag string is a compact symbolic description that could seed planning or skill libraries for deformable objects.","An inter-annotator agreement study on the same ten videos would quantify how much of the clustering benefit is intrinsic to the taxonomy rather than to the original labeler's judgment.","Because the taxonomy excludes tearing and plastic deformation, a direct extension would be to add irreversible deformation categories and check whether the action-ID clusters remain distinct when objects break or permanently yield."],"forward_implications":["Manipulation tasks can be described in a machine-readable tag form, so datasets and benchmarks for deformable-object manipulation can be compared on a common vocabulary.","Gripper and end-effector design can be driven by the deformation and interaction tags a task requires, e.g., preferring wide non-prehensile contacts over pinch grasps when object integrity matters.","Policies trained on one task may transfer to another task that shares the same T-DOM action-IDs, reducing retraining for new deformable-object skills.","The taxonomy separates actions that prior taxonomies merge, such as grasping a folded towel versus grasping a flat one, because the deformation state changes the control constraints.","Explicit tags for structured and unstructured bending can guide perception systems toward the relevant state features, such as loop counts or the number of accessible corners."],"supporting_citations":[{"why":"Baseline taxonomy for dexterous manipulation; supplies the contact, motion, and slippage categories and the action-segmentation style that T-DOM extends to deformable objects.","marker":"Bullock et al. (2012)"},{"why":"Baseline motion taxonomy that already distinguishes rigid and soft contacts and temporal versus permanent deformation; T-DOM compares its clusters against this taxonomy.","marker":"Paulius et al. (2020)"},{"why":"Classic grasp taxonomy that T-DOM contrasts with because it assumes rigid objects and classifies grasps by hand and object geometry.","marker":"Cutkosky (1989)"},{"why":"Grasp taxonomy categorizing human grasp types by number of fingers; used as a comparison point for prehensile grasp coverage.","marker":"Feix et al. (2015)"},{"why":"Provides the prehension-geometry viewpoint for cloth grasps that T-DOM adapts into point-constraint versus line-constraint grasp categories.","marker":"Borràs et al. (2020)"},{"why":"Introduces g-folds for folding cloth, which T-DOM adopts as the metric for structured bending in 2D objects.","marker":"Miller et al. (2012)"},{"why":"Supplies the definition of deformation as shape change from external force and surveys the sensing and measurement challenges for deformable objects.","marker":"Sanchez et al. (2018)"},{"why":"Standard materials-science source for the five deformation types—compression, tension, bending, torsion, shear—that structure the deformation axis.","marker":"Callister Jr and Rethwisch (2018)"},{"why":"Provides the dynamic versus quasi-static manipulation distinction that T-DOM recasts in terms of kinetic, elastic, and gravitational energy.","marker":"Mason (2001)"}],"fun_headline_variants":["Deformation labels split robot skills old taxonomies merge","New robot taxonomy counts deformations to tell skills apart","Deformation type now separates robot manipulation skills","Missing deformation data hides differences in robot skills","T-DOM taxonomy adds bending and torsion to robot skills"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison rests on one person manually deciding, by eye, when a robot action changes a taxonomy category, and no second labeler or independent deformation measurement is reported, so the claimed distinctions could partly reflect labelling choices rather than objective task structure.","fun_headline_variants_meta":{"raw":{"variants":["Deformation labels split robot skills old taxonomies merge","New robot taxonomy counts deformations to tell skills apart","Deformation type now separates robot manipulation skills","Missing deformation data hides differences in robot skills","T-DOM taxonomy adds bending and torsion to robot skills"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000825,"raw_usage":{"total_tokens":3636,"prompt_tokens":1005,"completion_tokens":2631,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":2557}},"tokens_in":621,"tokens_out":2631,"duration_ms":19233,"temperature":1.0,"reasoning_tokens":2557,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:04:27.667866+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two independent annotators label the same ten task videos with T-DOM. If they agree poorly on where action boundaries lie, or if an automated deformation measure (such as the derivative of the Gauss Linking Integral) fails to change exactly when the manual deformation labels change, the claim that deformation categories objectively differentiate the skills would be weakened.","supporting_citations":[{"cited_title":"IEEE transactions on Haptics 6(2): 129--144","cited_arxiv_id":null,"evidence_quote":"Baseline taxonomy for dexterous manipulation; supplies the contact, motion, and slippage categories and the action-segmentation style that T-DOM extends to deformable objects."},{"cited_title":"In: Proceedings of Robotics: Science and Systems","cited_arxiv_id":null,"evidence_quote":"Baseline motion taxonomy that already distinguishes rigid and soft contacts and temporal versus permanent deformation; T-DOM compares its clusters against this taxonomy."},{"cited_title":"IEEE Transactions on Robotics and Automation 5(3): 269--279","cited_arxiv_id":null,"evidence_quote":"Classic grasp taxonomy that T-DOM contrasts with because it assumes rigid objects and classifies grasps by hand and object geometry."},{"cited_title":"IEEE Transactions on human-machine systems 46(1): 66--77","cited_arxiv_id":null,"evidence_quote":"Grasp taxonomy categorizing human grasp types by number of fingers; used as a comparison point for prehensile grasp coverage."},{"cited_title":"The International Journal of Robotics Research 31(2): 249--267","cited_arxiv_id":null,"evidence_quote":"Introduces g-folds for folding cloth, which T-DOM adopts as the metric for structured bending in 2D objects."},{"cited_title":"The International Journal of Robotics Research 37(7): 688--716","cited_arxiv_id":null,"evidence_quote":"Supplies the definition of deformation as shape change from external force and surveys the sensing and measurement challenges for deformable objects."},{"cited_title":"10 edition","cited_arxiv_id":null,"evidence_quote":"Standard materials-science source for the five deformation types—compression, tension, bending, torsion, shear—that structure the deformation axis."},{"cited_title":"Cambridge, MA, USA: MIT Press","cited_arxiv_id":null,"evidence_quote":"Provides the dynamic versus quasi-static manipulation distinction that T-DOM recasts in terms of kinetic, elastic, and gravitational energy."}],"review_version":1}