{"id":"a0a50563-3f31-46d6-a580-d29477678bcd","arxiv_id":"2412.05066","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"BimArt synthesizes diverse, physically plausible bimanual hand motions for articulated objects from object trajectories alone, using a part-aware basis-point representation and predicted distance-based contact maps.","lead":"This paper introduces BimArt, a neural method that generates natural two-handed motions for manipulating objects with moving parts, like scissors or laptop lids, given only the object's trajectory. It uses learned contact maps as an intermediate step, avoiding the need for predefined grasps, and reports better contact quality than prior methods on two benchmarks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline claim 'surpasses the state of the art in quality and diversity' is not supported: CAMS-B beats BimArt on ARCTIC multimodality and acceleration, and on HOI4D scissors BimArt has the worst penetration; the user study excludes CAMS-B.","rationale":"The reader's scale-normalization concern (Eq. 1 and Supp. B) is reasonable but not the most load-bearing issue for the central claim: the heuristic articulation angle is chosen per object category and therefore does not invalidate the reported benchmark numbers, only the stronger generalization claim. The decisive gap is the mismatch between the headline 'surpasses the state of the art' and the paper's own tables, where CAMS-B wins the paper's diversity and smoothness metrics on ARCTIC and BimArt has the worst penetration on HOI4D scissors. The user study omits exactly that baseline, so the qualitative superiority claim is not tested against the strongest numeric competitor. The method has real strengths: the part-based BPS representation is novel and its ablation shows consistent gains; the optimization post-processing reduces ARCTIC penetration from 20.27% to 2.03% while preserving contact; the HOI4D cross-category comparison beats CAMS-X; and the reported nearest-neighbor distance (15.08 cm) argues against direct overfitting. These strengths justify a conditional acceptance, but the state-of-the-art claim needs revision, missing baseline comparisons, uncertainty estimates, and code release. My read therefore does not change the reader's verdict.","tokens_in":18437,"tokens_out":6528,"duration_ms":68532,"concrete_test":"Run the same 55-participant forced-choice user study with CAMS-B included as a third baseline, using the same 40-pair protocol and object coverage described in Sec. 4.2, and report per-metric 95% confidence intervals over at least 5 training seeds. If BimArt is not significantly preferred over CAMS-B at p<0.05, or if the confidence intervals for Mul/Accel exclude CAMS-B's point estimates, the state-of-the-art quality claim should be downgraded to 'competitive on contact/articulation, worse on diversity/smoothness.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that BimArt surpasses the state of the art in motion quality and diversity. The reported evidence does not establish this. On ARCTIC (Tab. 1), CAMS-B has higher multimodality (8.5602 vs 6.9093 cm) and lower acceleration (0.11959 vs 0.18846 cm/s^2), which are the paper's own diversity and smoothness metrics; BimArt's wins are penetration (2.03% vs 42.52%) and contact/articulation percentages. On HOI4D scissors (Tab. 2), BimArt w/ opt has penetration 0.464% versus CAMS 0.004% and Ours 0.591% versus 0.080%, so no category-specific baseline is outperformed on all relevant axes. The perceptual user study excludes CAMS-B, the baseline with the best diversity/smoothness numbers, and all metrics are single-seed point estimates with no confidence intervals. Thus the 'surpasses' claim depends on weighting penetration/contact above diversity/smoothness without a head-to-head perceptual justification and without uncertainty quantification. If users prefer CAMS-B or the intervals overlap, the claim collapses to 'competitive on contact metrics, worse on diversity/smoothness.'","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents BimArt, a three-stage generative pipeline for synthesizing bimanual hand motions interacting with articulated objects, given only the 7-DoF trajectory of the object. It introduces a canonicalized, part-based BPS object representation, a diffusion-based contact map generator, and a diffusion-based motion generator that uses contact maps as conditioning and as guidance, followed by an optimization-based refinement. The method is evaluated on ARCTIC and HOI4D, and the authors claim state-of-the-art motion quality and diversity, with an ablation study and a perceptual user study.","tokens_in":18734,"tokens_out":7688,"duration_ms":68898,"significance":"If the method's claims are substantiated, the contribution is meaningful: it removes the need for a reference grasp or coarse hand trajectory, supports simultaneous articulation and global motion, and uses a single cross-category model. The paper's strengths include a clear description of the representation and losses, a thorough ablation over representations and contact usage, and a sensitivity analysis in the supplement. However, the headline 'surpasses the state of the art in motion quality and diversity' is not fully supported by the reported tables, as discussed below. The method does show very low penetration on ARCTIC compared to adapted baselines, which is a valuable result.","major_comments":[{"comment":"The central claim that BimArt 'surpasses the state of the art in motion quality and diversity' is not supported by the reported evidence. On ARCTIC (Table 1), CAMS-B has higher multimodality (8.5602 vs 6.9093 cm) and lower acceleration (0.11959 vs 0.18846 cm/s^2), which are the paper's own diversity and smoothness metrics. On HOI4D scissors (Table 2), BimArt with and without optimization has higher penetration than CAMS (0.591% and 1.204% vs 0.080%). The perceptual user study (Sec. 4.2) excludes CAMS-B, the baseline that is strongest on these axes, so no head-to-head perceptual justification exists for weighting penetration/contact above diversity/smoothness. In addition, all quantitative results are single-seed point estimates without confidence intervals, so it is unclear whether any of the differences are statistically significant. Please either temper the claim to 'competitive or better on contact/articulation metrics with substantially lower penetration on ARCTIC' or add a comparison that includes CAMS-B in the user study and report variance across seeds.","section":"4.1, 4.2, Tables 1 and 2"},{"comment":"The 'category-agnostic' and 'unified' representation claim is weakened by the per-object-type heuristic used for the articulation angle in scale normalization: the supplement states that the angle is set to pi/2 for mixer and capsule machine, 0 for scissors and espresso machine, and pi for all others. This means that for a new object category, the user must supply a canonicalization rule, which is not category-agnostic. Since every downstream quantity (BPS features, contact maps, hand guidance) depends on this scale, the method's generalization to unseen articulated objects is at risk. Please provide a sensitivity analysis showing how the choice of this angle affects the final metrics, or revise the claim to specify that the canonicalization rule is object-model-specific.","section":"3.1 (Eq. 1) and Supplementary B"},{"comment":"The user study only compares BimArt against MDM-B and OMOMO-B, while the caption of Fig. 7 states that 'Our method outperforms the existing state of the art for all objects.' Because CAMS-B is excluded and MDM-B/OMOMO-B are adapted baselines, this statement overstates the scope of the perceptual evidence. Please either include CAMS-B in the user study or explicitly limit the claim to the compared baselines.","section":"4.2, Fig. 7"}],"minor_comments":[{"comment":"The sentence 'The single-hand variant for HOI4D dataset is denoted as MDM-U' appears to be a typo; Table 2 lists both MDM-U and OMOMO-U, so the text should refer to 'OMOMO-U' for the OMOMO variant.","section":"Section 4, baseline details"},{"comment":"The notation M(\\hat{X}(t), t, Z_o∅) is confusing; please define the null contact token and clarify which prediction is conditional and which is unconditional in the classifier-free guidance equations.","section":"Eq. (7)"},{"comment":"The CM metric is computed against the contact maps predicted by the contact model, which is also used to guide the motion model; please state explicitly in the table caption that this is an internal consistency metric and not comparable across methods that do not share the same contact predictor.","section":"Table 3 and Sec. 4.4"},{"comment":"The part index p in O^p_i is not defined before use; please state explicitly that p ∈ {top, bottom}.","section":"Section 3.1, Eq. (2)"},{"comment":"The concurrent work ManiDext [80] is mentioned as an exception but never compared quantitatively; please add a sentence explaining why it is excluded (e.g., different input assumptions or different evaluation protocol).","section":"Section 2, Related Work"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is well-structured and the method is technically solid, but the evidence does not yet support the 'state of the art' claim. I recommend requesting a revision that tempers the claims and adds the missing comparisons or uncertainty quantification. The supplementary sensitivity analysis is a positive sign."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: BimArt is a real step forward for bimanual hand-object synthesis with articulated objects, and the part-based BPS representation plus the distance-based contact maps are genuinely new. But the paper's headline claim that it surpasses the state of the art in quality and diversity is not actually supported by its own tables, because CAMS-B beats it on multimodality and acceleration on ARCTIC, and on HOI4D scissors its penetration is worse than several baselines. The user study also excludes CAMS-B.\n\nWhat's good: The method is well described, the ablations are thorough and support the design choices (part-based BPS, contact conditioning, guidance, optimizer). The cross-category unified model is a nice target, and the reported penetration rate on ARCTIC is dramatically lower than the adapted baselines. The supplementary sensitivity analysis for the post-processing weights is more than most papers do. The citation pattern looks fair; the novel pieces relative to CAMS, ManipNet, etc. are correctly identified.\n\nSoft spots: The 'surpasses SOTA' claim depends on weighting penetration/contact above diversity/smoothness, but there is no head-to-head perceptual comparison with CAMS-B, which has the best diversity and smoothness numbers. All metrics are single-seed point estimates with no confidence intervals. No code is released. The scale normalization relies on a heuristic articulation angle per object (Supp. B); if that angle is misestimated, the BPS features and generated contact maps shift. That's a modeling fragility worth acknowledging. The contact-map discrepancy metric (CM) is self-referential when used for ablation, though that's only an internal consistency check and the real metrics are against ground truth.\n\nBottom line: This is a competent, genuinely novel piece of work that deserves serious referee time. The revisions should be about calibration: soften the SOTA claim, add error bars, include CAMS-B in the user study, and ideally release code. For a reader in hand-object interaction or character animation, this is worth engaging with despite the overclaiming.","headline":"Genuinely new representation and solid ablations, but the SOTA claim outruns the numbers; worth reviewing with major revisions.","tokens_in":19330,"tokens_out":2435,"would_cite":true,"duration_ms":23610,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"BimArt generates realistic two-handed animations for articulated objects from only the object's 7-degree-of-freedom trajectory, without needing a reference grasp, a coarse hand trajectory, or separate grasping and articulating stages.","keywords":["bimanual interaction","articulated objects","hand motion synthesis","diffusion models","contact maps","basis point sets","MANO hand model","object manipulation"],"falsifier":"Take an articulated object not in the training categories whose maximum-extent articulation angle differs from the heuristic choices (0, $\\pi/2$, or $\\pi$), such as a folding chair or compound hinge, and run BimArt with its true trajectory: if penetration and contact or articulation percentages degrade sharply relative to the same object canonicalized with its correct angle, the heuristic is the load-bearing bottleneck. A sharper test is to feed a prismatic-joint object such as a sliding drawer, which violates the paper's two-part rotational-joint canonical frame (articulation axis aligned with the negative z-axis); the model should fail to produce stable contact on the moving part, delimiting the method's scope to rotational two-part articulated objects.","tokens_in":18209,"feed_emoji":"🤲","tokens_out":8493,"duration_ms":224747,"temperature":0.7,"pith_summary":"BimArt takes a single compact input—the trajectory of an articulated object, meaning its position, orientation, and articulation angle over time—and outputs a full two-handed animation of a person manipulating that object. The paper's central claim is that this can be done without any reference grasp, coarse hand trajectory, or separate grasping-versus-articulating stages, using a unified model trained across object categories. The method first generates distance-based contact maps that say which object regions each hand should touch at each frame, then uses those maps to guide a second generative model that produces hand surface keypoints and direction vectors to the object. On the ARCTIC dataset, BimArt reports a much lower hand-object penetration rate (2.03% of frames at the 1 cm threshold, versus roughly 30-67% for adapted baselines) while keeping contact and articulation percentages high, and a user study found respondents preferred its motions over two adapted baselines.","feed_headline":"Two-handed animation from just the object's path","feed_subtitle":"BimArt predicts touch points on the object, then animates the fingers; no reference grasp required.","key_machinery":"The load-bearing objects are the articulation-aware part-based BPS features, the bimanual contact maps, and the two diffusion models. Basis Point Sets (BPS) encode a geometry by storing, for each fixed basis point in space, the vector to the nearest object vertex; BimArt's extension computes these vectors separately for the top and bottom articulated parts after normalizing the object's scale, so that a bottle lid is sampled as densely as the bottle body. The contact maps are per-hand per-frame arrays that store, on each sampled object vertex, the minimum distance to any hand vertex; because they are distance-based and spatially embedded on the object, they reveal rich bimanual grasping patterns while leaving hand-object correspondence unspecified, which preserves diversity. The motion model generates hand surface keypoints plus direction vectors to the nearest object vertex, and uses the predicted contact maps in two ways: as a conditioning token with random dropout (classifier-free guidance) and as a gradient-based discrepancy target during denoising, pulling the generated hand geometry toward the predicted contact regions. The final MANO fitting optimization removes residual penetration and jitter with projection, penetration, and acceleration energy terms.","core_discovery":"On its own terms, the paper establishes that bimanual manipulation with articulated objects can be decomposed into contact-map prediction followed by contact-conditioned motion synthesis, and that this decomposition is what removes the need for grasp references. The contact generation network, a transformer-based denoising diffusion model, predicts per-frame distance-based contact maps for the left and right hands directly from an articulation-aware object encoding; the motion network then generates hand surface keypoints and object-direction vectors, conditioned on those maps with classifier-free guidance and an extra gradient term that aligns the generated contact geometry with the predicted maps at every denoising step. The object encoding that makes both stages category-agnostic is a normalized, part-based Basis Point Set representation in which the same basis points are mapped separately to each articulated part after scale normalization, so small moving parts receive equal sampling density to large static parts. A final optimization-based MANO fitting step with projection, penetration, and acceleration energies turns the generated keypoints into clean hand meshes. The paper demonstrates, via quantitative metrics and a forced-choice user study, that this pipeline produces motions judged more natural than adapted versions of several prior methods.","pith_inferences":["My inference: the contact-map-first idea should transfer beyond hands to whole-body interaction with articulated furniture or tools, where a sparse intent map on the object surface could mediate between environment geometry and full-body motion.","My inference: the scale-normalization heuristic (a per-object guess of the articulation angle that maximizes extent) is the most likely point of failure for genuinely new objects; replacing it with a data-driven canonicalization would probably improve zero-shot generalization more than any other component change.","My inference: the paper's multi-modality metric actually shows lower raw diversity than one adapted baseline (CAMS-B) on ARCTIC; the meaningful claim is diversity combined with physical plausibility, since that baseline's higher diversity comes with far more penetration."],"forward_implications":["Animators and VR developers can generate multiple plausible two-handed interactions from a single object trajectory, without authoring grasps by hand.","A single model handles multiple object categories and can execute object translation, rotation, and articulation simultaneously, which prior articulated-object methods could not do together.","Because contact maps are an explicit intermediate output, users or downstream systems could edit or re-target those maps to steer hand placement while keeping the rest of the motion generation intact.","The low penetration rate (2.03% of frames at 1 cm on ARCTIC) suggests the output motions can serve as initialization for physics-based simulators or motion retargeting pipelines without heavy cleanup."],"supporting_citations":[{"why":"Supplies the basis point set representation that BimArt extends with part-based, scale-normalized features.","marker":"[51]"},{"why":"Provides the ARCTIC dataset of bimanual interactions used as the primary training and evaluation data.","marker":"[14]"},{"why":"Provides the HOI4D dataset and evaluation protocol used for cross-category comparison.","marker":"[40]"},{"why":"Defines the MANO hand model that parameterizes and fits the generated hand motions.","marker":"[52]"},{"why":"Introduces classifier-free guidance, used to condition the motion model on contact features.","marker":"[23]"},{"why":"Defines the DDPM noise schedule used to train both diffusion models.","marker":"[24]"},{"why":"CAMS, the main prior method for articulated manipulation, which BimArt adapts into a bimanual baseline and outperforms.","marker":"[83]"},{"why":"MDM, the diffusion motion model adapted as the MDM-B baseline for comparison.","marker":"[62]"}],"fun_headline_variants":["No reference grasp: BimArt animates two hands automatically","BimArt: from object motion to bimanual interaction, no grasp refs","Contact maps unlock BimArt's unified hand-object animation","Object path alone drives BimArt's bimanual synthesis","BimArt: unified generation of bimanual interaction, no grasp needed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method's object encoding relies on a per-category guess about which articulation angle makes each object span its largest extent when normalizing scale; if that guess is wrong for an unfamiliar object, every downstream quantity—the BPS features, the contact maps, and the hand guidance—is computed from a distorted geometry, and the model's generalization degrades.","fun_headline_variants_meta":{"raw":{"variants":["No reference grasp: BimArt animates two hands automatically","BimArt: from object motion to bimanual interaction, no grasp refs","Contact maps unlock BimArt's unified hand-object animation","Object path alone drives BimArt's bimanual synthesis","BimArt: unified generation of bimanual interaction, no grasp needed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00072,"raw_usage":{"total_tokens":3233,"prompt_tokens":944,"completion_tokens":2289,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":560,"completion_tokens_details":{"reasoning_tokens":2196}},"tokens_in":560,"tokens_out":2289,"duration_ms":16994,"temperature":1.0,"reasoning_tokens":2196,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:56:25.415264+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take an articulated object not in the training categories whose maximum-extent articulation angle differs from the heuristic choices (0, $\\pi/2$, or $\\pi$), such as a folding chair or compound hinge, and run BimArt with its true trajectory: if penetration and contact or articulation percentages degrade sharply relative to the same object canonicalized with its correct angle, the heuristic is the load-bearing bottleneck. A sharper test is to feed a prismatic-joint object such as a sliding drawer, which violates the paper's two-part rotational-joint canonical frame (articulation axis aligned with the negative z-axis); the model should fail to produce stable contact on the moving part, delimiting the method's scope to rotational two-part articulated objects.","supporting_citations":[{"cited_title":"Ef- ficient learning on point clouds with basis point sets","cited_arxiv_id":null,"evidence_quote":"Supplies the basis point set representation that BimArt extends with part-based, scale-normalized features."},{"cited_title":"Black, and Ot- mar Hilliges","cited_arxiv_id":null,"evidence_quote":"Provides the ARCTIC dataset of bimanual interactions used as the primary training and evaluation data."},{"cited_title":"Hoi4d: A 4d egocentric dataset for category-level human- object interaction","cited_arxiv_id":null,"evidence_quote":"Provides the HOI4D dataset and evaluation protocol used for cross-category comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the MANO hand model that parameterizes and fits the generated hand motions."},{"cited_title":"Classifier-free diffusion guidance","cited_arxiv_id":null,"evidence_quote":"Introduces classifier-free guidance, used to condition the motion model on contact features."},{"cited_title":"Denoising dif- fusion probabilistic models","cited_arxiv_id":null,"evidence_quote":"Defines the DDPM noise schedule used to train both diffusion models."},{"cited_title":"Cams: Canonicalized manipulation spaces for category-level functional hand-object manipulation synthe- sis","cited_arxiv_id":null,"evidence_quote":"CAMS, the main prior method for articulated manipulation, which BimArt adapts into a bimanual baseline and outperforms."},{"cited_title":"Human motion diffu- sion model","cited_arxiv_id":null,"evidence_quote":"MDM, the diffusion motion model adapted as the MDM-B baseline for comparison."}],"review_version":1}