{"id":"67efaabd-79a5-43db-93e1-e1d38c89f880","arxiv_id":"2607.11688","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A parametric linkage face template plus hierarchical collision-driven optimization synthesizes manufacturable facial mechanisms from 2D portraits and runs them with dual-identity conversational motion.","lead":"The paper builds an automated pipeline that turns a single 2D portrait into a collision-free, 3D-printable linkage-driven animatronic face, then drives it with dual-speaker conversational motion for speaking and listening. If it scales, personalized social-robot faces stop being one-off engineering projects.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The 66.7% success rate and fixed-topology failures undercut the claim of automated synthesis for a wide range of facial morphologies.","rationale":"The Reader correctly isolates the fixed-topology assumption as the structural weak point. Table I and App. D7 already show that 5/15 heads leave the constraint set empty under continuous parameter tuning alone; that is not an optimization detail but a limit on the automation claim for “a wide range of facial morphologies.” The hierarchical pipeline is still a genuine engineering advance over the stated baselines and is demonstrated on real hardware and diverse characters, so REJECT is too strong. Unconditional ACCEPT would require either higher SR under the same template or an explicit topology/reconfigurability extension. Hence CONDITIONAL is the right verdict; the concrete test above would decide whether the condition can be relaxed with modest changes or remains a hard bound. No stronger internal inconsistency appears; the dual-identity motion and mapping results are secondary to the hardware-synthesis claim.","tokens_in":38991,"tokens_out":574,"duration_ms":5312,"concrete_test":"Re-run Algorithm 1 on the five failed heads after (a) allowing discrete topology variants (e.g., drop one mouth-corner DoF or switch lower-lip to a simpler slider) or (b) adding a single mechanical reconfigurability DoF (interpupillary or jaw-base slide). If SR remains ≤66.7% or still requires hand redesign, the fixed-template automation claim is limited to a narrower morphology class than advertised; if SR rises to ≥90% with only those discrete choices, the claim strengthens under a modestly expanded search.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim rests on hierarchical optimization of a fixed parametric template (six-bar brows, four-bar lids/lips/jaw, five-bar corners; eyes/jaw fixed after coarse init) producing collision-free, manufacturable mechanisms more successfully than baselines (Table I: 66.7% SR / 10 of 15). The paper itself documents that this topology is structurally infeasible for some morphologies: pointed chins with insufficient internal clearance, and multi-module interlocks that pose updates + amplitude decay (Alg. 1, Eq. 2, γ-scheduling) cannot resolve (App. D7). Those 5/15 failures are not mere optimization misses; they are empty feasible sets under the fixed topology. If the target population includes such geometries (or if the 15-head set is not representative), the automation claim does not hold without topology search or reconfigurable hardware. Manual comparison is only one compact head; no open code/data lets outsiders re-run the 15 cases.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper presents an end-to-end pipeline for automated synthesis of linkage-driven animatronic faces from a single 2D portrait, together with a dual-identity conversational motion system for physical robots. A fixed parametric template (eyebrow six-bar, eye/eyelid four-bars, lip four-bars, mouth-corner five-bars, jaw four-bar, 3-DoF neck) is adapted by hierarchical optimization: an inner loop maximizes AU-derived trajectory amplitudes under anatomy-guided feasible volumes (Eq. 1), and an outer loop resolves interferences via MTV/QP base-pose updates or amplitude scheduling (Alg. 1, Eq. 2). Separately, a turn-aware dual-audio Transformer predicts blendshapes for speaking and listening, mapped region-wise to motors. Experiments report 66.7% success on 15 heads (Table I), one manual-design comparison, motion/mapping benchmarks (Tables II–III), physical demos, and N=100 user studies.","tokens_in":39334,"tokens_out":1450,"duration_ms":18467,"significance":"If the automation results hold under a clearly scoped morphology class, the work addresses a genuine bottleneck: bespoke mechanical redesign that currently prevents large-scale personalization of animatronic faces. The hierarchical split of kinematic synthesis and collision refinement is a practical engineering contribution, and the dual-identity real-time speaking/listening stack with physical deployment and perceptual studies is a meaningful step beyond monologue talking heads. Strengths include physical builds, explicit ablations of the design loop (Table I), comparison to manual design on a compact head, on-device FPS, and a reasonably sized user study. The fixed-topology premise and incomplete success rate limit how far the “wide range / scalable personalization” claim currently reaches, but the systems contribution remains substantial for robotics and HRI.","major_comments":[{"comment":"Abstract and §I claim automated synthesis for a “wide range of facial morphologies,” but Table I reports only 66.7% success (10/15), and App. D7 attributes the remaining failures to empty feasible sets under the fixed topology (insufficient internal clearance; multi-module interlocks that pose updates and γ-scheduling cannot resolve). Those are structural template failures, not mere optimizer misses. The central automation claim therefore needs either (i) a precise characterization of the morphology class for which the template is feasible, with failure rates broken down by geometry type, or (ii) an explicit limitation that topology search / reconfigurable hardware is required outside that class. Without this, “wide range” overstates what Alg. 1 delivers.","section":null},{"comment":"§V-B, “Comparative efficiency against manual design”: the manual baseline is two senior designers on a single compact head (β=0.84), 22.8 h vs 11.7 min. This is useful as a case study but is too thin to support a general claim of superiority over manual mechanical design. At minimum, report designer protocols, success criteria, and preferably additional heads (including a failure-mode geometry from App. D7). Otherwise reframe the claim as a single-instance efficiency illustration rather than a comparative evaluation pillar.","section":null},{"comment":"Table I: Global Joint-Opt has 0% SR and very large IV. That baseline is informative only if it is a fair formulation of joint optimization (soft collision penalties, same L-BFGS budget, same initialization). Please specify the full joint objective weights (App. D5), iteration budgets, and whether any multi-start or constraint-handling variants were tried. If the joint baseline is essentially unsolvable as posed, the hierarchical gain is partly architectural necessity rather than empirical dominance; that should be stated carefully so the 66.7% figure is not over-interpreted.","section":null},{"comment":"§III-B and App. D6–D7: the 15-head evaluation set is central to the hardware claim, but the paper does not fully specify selection criteria, scale distribution, or how many of the eight visualized identities in Fig. A.5 are among the 10 successes. Please tabulate per-head outcomes (success/fail, final Exp, IV, outer-loop iterations) and relate failures to measurable geometric features (e.g., chin clearance, inter-module volume). Without that, reproducibility and external validity of the 66.7% rate remain limited.","section":null}],"minor_comments":[{"comment":"Eq. (1): L_amp = 1/α_k with a maximization intent via smaller L is clear, but the multi-trajectory case (K trajectories, shared or per-k Φ) should state whether Φ is shared across trajectories of the same module and how α_k are aggregated into the reported Exp = Σ α_k.","section":null},{"comment":"Fig. 3 / Alg. 1: clarify whether eyes and jaw are ever re-optimized after coarse init when outer-loop collisions involve those modules, or whether only brow/mouth bases move.","section":null},{"comment":"Table II: DualTalk is offline (uses partner future motion); mark this more prominently in the table caption so the RT column is not the only cue.","section":null},{"comment":"Table IV total (25.84 h) assumes parallel Stages 2–3; state single-printer sequential time as well for readers estimating lab throughput.","section":null},{"comment":"App. C Morpheus comparison (13.3% SR) is valuable; move a short summary into the main text near Table I so architectural sensitivity is not buried.","section":null},{"comment":"Notation: p_base vs p^base, and α_limit vs α^limit, are used inconsistently between main text and Algorithm 1; unify.","section":null},{"comment":"User study (Fig. 7): report statistical tests (e.g., paired tests or CI) for Ours vs ablations rather than means alone.","section":null},{"comment":"Related work: briefly position against other parametric/retargetable facial robots beyond Morpheus to clarify novelty of the hierarchical collision loop.","section":null}],"recommendation":"major_revision","confidential_remarks":"Solid systems paper with real hardware and a coherent end-to-end story; appropriate for a strong robotics venue after the automation claim is scoped to the fixed template’s feasible set. The fixed-topology limitation is the main load-bearing issue, not circularity or fraud. Overlap with the authors’ Morpheus line is natural but should stay clearly differentiated (automation + dual-identity conversation vs hybrid actuation expressiveness). I would not reject on novelty grounds if the revision tightens claims and reports per-head outcomes."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is that they ship a working hierarchical pipeline: fixed parametric linkage template (six-bar brows, four/five-bar mouth pieces, etc.), anatomy/AU-guided inner kinematic opt, MTV/QP outer collision loop, then dual-audio speaking+listening motion mapped region-wise to motors. It is not pure novelty in any single piece—Coros/Thomaszewski-style linkage synthesis, DualTalk-style dyadic heads, and Morpheus/EMO-style faces are all cited—but the combination is concrete and runs on hardware for diverse characters (Yoda, Jack, Chinese-folktale heads, etc.).\n\nWhat they do well: Table I is clean. Hierarchical design hits 10/15 success vs 0–33% for global joint-opt, local-only, and local+heuristic; expressiveness stays comparable while collision volume drops. Manual comparison on one compact head (two designers, 22.8 h vs ~12 min) shows tighter upward corner trajectories. Motion tables beat speaker-only and DualTalk on their blendshape dataset; mapping beats landmark/FLAME/global-blendshape baselines on LSE and FPS. N=100 user study favors the full system over the obvious ablations. Kinematic appendices are thorough; free parameters are listed and used consistently.\n\nSoft spots, in proportion: the fixed topology is the structural limit they themselves flag in App. D7. Five of fifteen heads fail because the feasible set is empty (pointed chin clearance, multi-module interlocks that pose updates + γ-scheduling cannot fix). That undercuts the “wide range of morphologies” language; automation here is parameter retargeting, not topology search. Manual baseline is only one head. No code or data release, so outsiders cannot re-run the 15 cases. Those are real but not load-bearing flaws that make the central claim false—the pipeline demonstrably produces manufacturable, conversational faces faster than the stated alternatives when the template fits.\n\nThis is for HRI / entertainment-robotics people who care about personalization bottlenecks and physical execution, not for pure mechanism theorists. Math and citation pattern look solid; data are empirical and multi-part. I would send it to peer review; the contribution is accept-shaped for cs.RO once success-rate limits and artifacts are handled. Worth engaging if you work on scalable social robots.","headline":"Solid systems paper that actually automates high-DoF face mechanisms from a portrait and drives them in dual-turn conversation; fixed-topology failures are real but already documented and do not erase the demonstrated pipeline.","tokens_in":39961,"tokens_out":595,"would_cite":true,"duration_ms":7157,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A fixed parametric linkage template plus hierarchical optimization can turn one 2D portrait into a manufacturable, collision-free animatronic face that supports real-time speaking and listening conversation.","keywords":["animatronic face","automatic mechanism design","linkage synthesis","hierarchical optimization","conversational motion","speaking and listening","expression mapping","social robots"],"falsifier":"Run the pipeline on a larger, more extreme set of head geometries (for example, very pointed chins or non-human proportions outside the reported 15-head suite) and measure whether the success rate remains near the claimed two-thirds without any topology change.","tokens_in":39861,"feed_emoji":"🤖","tokens_out":680,"duration_ms":5522,"temperature":0.7,"pith_summary":"Today's high-fidelity robot faces are handcrafted for one head geometry at a time, so personalizing hundreds of distinct faces is too slow and expensive. This paper claims that automation is possible if one starts from a single modular, linkage-driven template whose topology and actuator layout are fixed but whose continuous parameters can be scaled and retargeted. From a single 2D portrait the system reconstructs a 3D mesh, initializes module base poses, then runs an inner loop that maximizes Action-Unit trajectories inside anatomy-guided feasible volumes and an outer loop that resolves collisions by minimum-translation pose updates or amplitude decay. The resulting designs are reported to succeed more often and faster than pure local optimization, global joint optimization, or heuristic repulsion, and to beat expert manual design on a compact head in both time and mouth-corner expressiveness. Separately, a dual-audio transformer that gates speaking versus listening produces blendshapes and head pose that map region-wise onto the physical motors at real-time rates, so the finished head can hold multi-turn dialogue rather than monologue. If the claim holds, personalized conversational robots become a manufacturing problem rather than a one-off craft project.","feed_headline":"One portrait becomes a talking robot face automatically","feed_subtitle":"Fixed linkage template plus collision-aware optimization yields manufacturable heads that speak and listen in real time.","key_machinery":"Hierarchical automatic design on a fixed modular template: anatomy-guided feasible volumes and AU-trajectory scaling inside each module, followed by a collision-driven outer loop that applies minimum-translation-vector pose updates or amplitude scheduling until the global CAD assembly is interference-free.","core_discovery":"A hierarchical automatic design algorithm built on one parametric linkage face template can take a single 2D portrait, reconstruct the 3D geometry, and synthesize a collision-free, manufacturable internal mechanism that is more successful and faster than local-only, global joint, or local-plus-heuristic baselines, while a dual-identity audio model supplies real-time speaking and listening motion suitable for physical execution.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Single portrait to collision-free robot face via hierarchical synthesis","Parametric linkages auto-build manufacturable animatronic faces from photos","One 2D portrait yields speaking-listening robot head mechanism","Algorithm synthesizes diverse facial mechanisms faster than manual design","Hierarchical design turns portraits into real-time conversational robot faces"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"A single fixed mechanism topology remains feasible for a wide range of facial shapes; when internal clearance is too small or modules interlock, the continuous-parameter optimizer cannot recover a valid design.","fun_headline_variants_meta":{"raw":{"variants":["Single portrait to collision-free robot face via hierarchical synthesis","Parametric linkages auto-build manufacturable animatronic faces from photos","One 2D portrait yields speaking-listening robot head mechanism","Algorithm synthesizes diverse facial mechanisms faster than manual design","Hierarchical design turns portraits into real-time conversational robot faces"]},"model":"grok-4.5","effort":"low","cost_usd":0.005168,"raw_usage":{"total_tokens":1494,"prompt_tokens":853,"num_sources_used":0,"completion_tokens":86,"cost_in_usd_ticks":51680000,"prompt_tokens_details":{"text_tokens":853,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":555,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":853,"tokens_out":86,"duration_ms":4609,"temperature":1.0,"reasoning_tokens":555,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T03:50:58.761819+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the pipeline on a larger, more extreme set of head geometries (for example, very pointed chins or non-human proportions outside the reported 15-head suite) and measure whether the success rate remains near the claimed two-thirds without any topology change.","supporting_citations":[],"review_version":1}