{"id":"84c8e6b2-027e-4448-a542-3a6fc4d480f2","arxiv_id":"2607.07880","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A text-conditioned diffusion model using dynamic object-centric BPS, mixed-domain training, and contact augmentation produces generalizable full-body locomotion-to-articulated-object interaction sequences that beat adapted baselines on contact and pose metrics.","lead":"GIRAF generates full-body human motions that walk up to and manipulate articulated objects such as drawers and refrigerators from text instructions and initial poses. It supplies synthetic interaction data useful for robotics training and virtual agents that must handle real-world object variation.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Generalization claim rests on reconstruction metrics and qualitative figures over a narrow, augmented ParaHome subset rather than held-out placements or novel articulations.","rationale":"The reader correctly flags the contact-based augmentation (§4.3) as a soft spot and issues a CONDITIONAL verdict pending code release and fuller ablations. That concern is real, yet the more load-bearing issue for the paper’s headline claim is the mismatch between the evaluation protocol (reconstruction on same-category sequences) and the language of \"strong generalization to unseen object configurations.\" The augmentation may itself introduce systematic kinematic artifacts that the later DDIM noise optimization merely masks; without a true held-out quantitative protocol the superiority numbers cannot be read as evidence of generalization. Because the engineering contribution remains solid and the qualitative examples are encouraging, the appropriate verdict stays CONDITIONAL rather than REJECT; the condition should simply be tightened to require a proper OOD split and ablations that isolate the augmentation and optimization stages. Agreement with the reader is therefore partial: same overall verdict, different primary soft spot.","tokens_in":14785,"tokens_out":546,"duration_ms":6868,"concrete_test":"Hold out an entire object category (or a spatial grid cell never seen in the 0.1 m augmentation of §4.3) and recompute Table 1 contact distance / F1 and Table 2 MPJPE / hand error on that held-out set without the scene-aware noise optimization of §4.4. If the gap over LINGO/CHOIS shrinks below 10 % or absolute contact distance exceeds 3 cm, the generalization claim is overstated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim asserts \"strong generalization to unseen object configurations\" that surpasses SOTA. Yet Tables 1–3 report reconstruction-style errors (MPJPE, contact distance, FID, etc.) conditioned on the exact initial pose, object state and text of test sequences drawn from the same four ParaHome categories (drawer, microwave, refrigerator, washing machine). Section 5.1 states that after contact-based augmentation the training set contains only 2100 sequences / 3 h; no quantitative split isolating truly novel placements, scales or joint types is given. Figures 4–6 supply qualitative evidence of closet doors and height variation, but these remain visual anecdotes. The conclusion itself concedes failure on unseen rotational mechanisms. Consequently the quantitative superiority may largely reflect better in-distribution fitting plus post-hoc noise optimization rather than the claimed out-of-distribution generalization.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"GIRAF proposes a text-conditioned diffusion model for synthesizing long-horizon full-body human motion that approaches, contacts, and actuates two-part articulated household objects (drawer, microwave, refrigerator, washing machine). The method combines three components: a dynamic object-centric basis-point-set (BPS) representation that jointly encodes object surface distances, hand end-effector distances and binary contact labels; a mixed-domain training strategy that uses FiLM conditioning and an annealing schedule over homogeneous locomotion/interaction batches; and a contact-preserving augmentation that relocates and rescales objects on a discrete grid while re-solving human motion via CCD inverse kinematics. After generation, a two-stage scene-aware DDIM noise optimization reduces contact error, penetration and jitter. Experiments on an augmented ParaHome+Babel corpus (2100 sequences, ~3 h) report consistent gains over retrained LINGO and CHOIS baselines on contact distance/F1, penetration, MPJPE/hand/object error, R-precision and FID, together with qualitative examples of height variation, multi-door closets and left/right/both-hand strategies.","tokens_in":15042,"tokens_out":817,"duration_ms":8059,"significance":"If the claimed generalization holds, the work would supply a practical, text-driven generator of coordinated locomotion-plus-articulated-manipulation sequences that current HOI and hand-only pipelines do not jointly address. The object-centric BPS encoding and the explicit mixed-domain schedule are concrete, reusable design choices for the growing literature on diffusion-based human-scene interaction. The paper is purely empirical; it does not ship code, proofs or parameter-free derivations, but the quantitative tables and qualitative figures already demonstrate measurable improvements on standard contact, reconstruction and text-to-motion metrics within the four ParaHome categories.","major_comments":[{"comment":"The central claim of 'strong generalization to unseen object configurations' (abstract, §1, §5.6) is not supported by a quantitative held-out protocol. Tables 1–3 report reconstruction-style metrics conditioned on the exact initial pose, object state and text of test sequences drawn from the same four ParaHome categories used for training. Section 5.1 states that after contact-based augmentation the corpus contains only 2100 sequences; no split isolating novel placements, scales or joint types is reported, and the conclusion itself concedes failure on unseen rotational mechanisms. Figures 4–6 supply only qualitative anecdotes. Without a quantitative OOD evaluation (e.g., held-out grid cells, novel object scales, or a fifth articulated category), the numerical superiority may largely reflect better in-distribution fitting plus post-hoc noise optimization rather than the claimed generaliza","section":null},{"comment":"Section 4.3’s contact-based augmentation (0.1 m grid, random rescaling, CCD-IK with rotational limits) is load-bearing for the diversity claim, yet no ablation quantifies residual kinematic artifacts or their effect on the diffusion prior. Because the later noise-optimization stage (§4.4) can mask many of these artifacts, it remains unclear how much of the reported contact and penetration gains (Table 1) are attributable to the learned model versus the post-processing. An ablation that reports Tables 1–2 both with and without the augmentation (and with/without noise optimization) is required to substantiate the contribution.","section":null},{"comment":"The mixed-domain training claim (§4.2) that 'homogeneous batches are imperative' is asserted without external controls. The paper reports neither an ablation that replaces homogeneous batches by mixed batches of equal size nor a comparison against a pure interaction-only baseline that still receives the locomotion mask and FiLM embeddings. Given that the annealing schedule and the 0.5 m locomotion threshold are free parameters, the necessity of the proposed schedule remains unproven.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is that GIRAF is a practical pipeline that actually produces coordinated approach-plus-manipulation sequences for drawers, microwaves, fridges and washers, with better contact distance, F1 and penetration than the two baselines they retrained with hands. That is useful for synthetic data pipelines.\n\nWhat is new is the object-centric dynamic BPS that folds surface distances, hand end-effectors and binary contact labels into one shared frame, plus the homogeneous-batch annealing schedule and the contact-preserving relocate+CCD-IK augmentation. These are known pieces (BPS, FiLM, DDIM noise opt) assembled into a coherent text-conditioned diffusion model that does not assume the human is already next to the object. Tables 1-3 show consistent gains on contact, MPJPE/hand/object error, R-precision and FID; the qualitative figures of height variation and closet doors look plausible. Related-work coverage is solid and the conclusion is honest about remaining foot-slide and novel rotational joints.\n\nThe soft spots are real but proportionate. The strongest claim of “strong generalization to unseen object configurations” rests on reconstruction metrics conditioned on the exact test initial pose/state/text from the same four ParaHome categories, plus a few visual anecdotes after heavy 0.1 m grid + rescale augmentation. No quantitative held-out placement split, no error bars, free parameters (0.5 m threshold, grid, K, annealing) under-ablated, and only ~3 h of data. The post-hoc noise optimization is doing real work. None of this is fatal for an engineering paper; it just means the generalization language should be dialed back.\n\nMath is standard diffusion, data and citations look clean. This is for people building HOI synthetic data or text-driven character animation. It deserves a serious referee. I would engage with the BPS idea and the mixed-domain schedule if the code appears; otherwise it is a solid incremental systems result worth knowing about.","headline":"Clean systems paper that finally joins full-body locomotion with two-DoF articulated manipulation under text; contact numbers improve, but the generalization claim outruns the narrow reconstruction eval.","tokens_in":15645,"tokens_out":500,"would_cite":true,"duration_ms":20100,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A text-conditioned diffusion model generates full-body human motion that approaches, contacts, and actuates articulated objects, generalizing to placements never seen in training.","keywords":["human-object interaction","articulated objects","motion diffusion","full-body motion synthesis","text-conditioned generation","contact modeling","locomotion-to-manipulation"],"falsifier":"Evaluate the trained model on object placements far outside the 0.1 m augmentation grid and on articulated mechanisms absent from training (for example multi-link doors); if contact distance or penetration then exceeds the baselines reported for the original test set, the generalization claim is falsified.","tokens_in":15686,"feed_emoji":"🚪","tokens_out":826,"duration_ms":20104,"temperature":0.7,"pith_summary":"Prior work either handles simple full-body actions with static objects or restricts itself to hand-only grasping, leaving open the harder problem of coordinated sequences that walk up to an articulated object, make fine contact, and move its parts. This paper claims that one diffusion model can close that gap when three ingredients are combined: an object-centric representation that places contact labels, hand end-effectors, and object surfaces in the same shared space; mixed-domain training that balances free locomotion with interaction; and contact-preserving relocation of training examples to expand spatial diversity. The resulting sequences are physically plausible over long horizons and adapt body reach, grasp, and object articulation to new positions and scales. If the claim holds, synthetic motion of this kind becomes a practical source of training data for robots and virtual agents that must navigate and manipulate the everyday world.","feed_headline":"Diffusion model walks up and opens drawers it never saw","feed_subtitle":"Shared contact points plus mixed training let full-body motion generalize across placements and shapes.","key_machinery":"Dynamic basis point sets (dynamic BPS) canonicalized to the articulated object part: a fixed cloud of points that jointly encode object-surface distances, distances to hand end-effectors, and binary contact labels, rendering contact shape-agnostic and transferable across geometries.","core_discovery":"Unifying hand–object contact, hand end-effectors, and object surfaces inside a dynamic object-centric basis-point representation, trained with mixed locomotion–interaction batches and contact-based data relocation, lets a single text-conditioned diffusion model synthesize seamless full-body sequences that approach, grasp, and actuate articulated objects while generalizing to unseen configurations and outperforming prior methods on contact, penetration, pose, and text-alignment metrics.","pith_inferences":["The same shared-point voting scheme for contact could extend to multi-object clutter or continuous multi-DoF articulations beyond the two-DoF hinges used here.","Coupling the diffusion prior more tightly with a physics simulator might remove residual foot sliding and joint jitter without a separate noise-optimization stage.","Text control over grasp laterality suggests a route to interactive virtual agents whose contact preferences can be specified on the fly."],"forward_implications":["Long-horizon locomotion-to-manipulation can be generated by one model without separate navigation and grasp stages.","Scarce paired human–scene data can be expanded by contact-preserving object relocation rather than new motion capture.","Text prompts can steer not only the action but also contact strategy (one hand versus both).","Synthetic sequences become usable training data for embodied agents that must both navigate and actuate everyday articulated objects."],"fun_headline_variants":["Object-centric diffusion opens unseen drawers with full-body motion","Text model walks grasps and actuates articulated objects never trained on","Unified contacts let diffusion generalize full-body drawer interactions","Mixed training synthesizes seamless approach and grasp of novel objects","Contact relocation trains one model for unseen object placements and shapes"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That relocating contacts onto a coarse 3-D grid and re-solving human pose with inverse kinematics plus joint limits still yields training examples realistic enough for the model to generalize to truly novel placements and scales.","fun_headline_variants_meta":{"raw":{"variants":["Object-centric diffusion opens unseen drawers with full-body motion","Text model walks grasps and actuates articulated objects never trained on","Unified contacts let diffusion generalize full-body drawer interactions","Mixed training synthesizes seamless approach and grasp of novel objects","Contact relocation trains one model for unseen object placements and shapes"]},"model":"grok-4.5","effort":"low","cost_usd":0.004366,"raw_usage":{"total_tokens":1305,"prompt_tokens":770,"num_sources_used":0,"completion_tokens":85,"cost_in_usd_ticks":43660000,"prompt_tokens_details":{"text_tokens":770,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":450,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":770,"tokens_out":85,"duration_ms":6618,"temperature":1.0,"reasoning_tokens":450,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T16:06:45.562075+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Evaluate the trained model on object placements far outside the 0.1 m augmentation grid and on articulated mechanisms absent from training (for example multi-link doors); if contact distance or penetration then exceeds the baselines reported for the original test set, the generalization claim is falsified.","supporting_citations":[],"review_version":1}