Pith. sign in

REVIEW 4 major objections 8 minor 155 references

Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots

T0 review · 4 major / 8 minor · reviewed 2026-07-14 · grok-4.5

Pith's one-line read A fixed parametric linkage template plus hierarchical optimization can turn one 2D portrait into a manufacturable, collision-free animatronic face that supports real-time speaking and listening conversation.

desk verdict Solid systems paper that actually automates high-DoF face mechanisms from a portrait and drives them in dual-turn conversation; fixed-topology failures are real but already documented and do not erase the demonstrated pipeline. read the letter →

arxiv 2607.11688 v1 pith:IJLQ7J33 submitted 2026-07-13 cs.RO

classification cs.RO
keywords animatronicfaceautomaticmechanismdesignlinkagesynthesishierarchicaloptimizationconversationalmotionspeakingandlisteningexpressionmappingsocialrobots
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Today's high-fidelity robot faces are handcrafted for one head geometry at a time, so personalizing hundreds of distinct faces is too slow and expensive. This paper claims that automation is possible if one starts from a single modular, linkage-driven template whose topology and actuator layout are fixed but whose continuous parameters can be scaled and retargeted. From a single 2D portrait the system reconstructs a 3D mesh, initializes module base poses, then runs an inner loop that maximizes Action-Unit trajectories inside anatomy-guided feasible volumes and an outer loop that resolves collisions by minimum-translation pose updates or amplitude decay. The resulting designs are reported to succeed more often and faster than pure local optimization, global joint optimization, or heuristic repulsion, and to beat expert manual design on a compact head in both time and mouth-corner expressiveness. Separately, a dual-audio transformer that gates speaking versus listening produces blendshapes and head pose that map region-wise onto the physical motors at real-time rates, so the finished head can hold multi-turn dialogue rather than monologue. If the claim holds, personalized conversational robots become a manufacturing problem rather than a one-off craft project.

What carries the argument

Hierarchical automatic design on a fixed modular template: anatomy-guided feasible volumes and AU-trajectory scaling inside each module, followed by a collision-driven outer loop that applies minimum-translation-vector pose updates or amplitude scheduling until the global CAD assembly is interference-free.

What would settle it

Run the pipeline on a larger, more extreme set of head geometries (for example, very pointed chins or non-human proportions outside the reported 15-head suite) and measure whether the success rate remains near the claimed two-thirds without any topology change.

Watch

Extended reading notes

Core claim

A hierarchical automatic design algorithm built on one parametric linkage face template can take a single 2D portrait, reconstruct the 3D geometry, and synthesize a collision-free, manufacturable internal mechanism that is more successful and faster than local-only, global joint, or local-plus-heuristic baselines, while a dual-identity audio model supplies real-time speaking and listening motion suitable for physical execution.

Load-bearing premise

A single fixed mechanism topology remains feasible for a wide range of facial shapes; when internal clearance is too small or modules interlock, the continuous-parameter optimizer cannot recover a valid design.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 8 minor

Summary. The paper presents an end-to-end pipeline for automated synthesis of linkage-driven animatronic faces from a single 2D portrait, together with a dual-identity conversational motion system for physical robots. A fixed parametric template (eyebrow six-bar, eye/eyelid four-bars, lip four-bars, mouth-corner five-bars, jaw four-bar, 3-DoF neck) is adapted by hierarchical optimization: an inner loop maximizes AU-derived trajectory amplitudes under anatomy-guided feasible volumes (Eq. 1), and an outer loop resolves interferences via MTV/QP base-pose updates or amplitude scheduling (Alg. 1, Eq. 2). Separately, a turn-aware dual-audio Transformer predicts blendshapes for speaking and listening, mapped region-wise to motors. Experiments report 66.7% success on 15 heads (Table I), one manual-design comparison, motion/mapping benchmarks (Tables II–III), physical demos, and N=100 user studies.

Significance. If the automation results hold under a clearly scoped morphology class, the work addresses a genuine bottleneck: bespoke mechanical redesign that currently prevents large-scale personalization of animatronic faces. The hierarchical split of kinematic synthesis and collision refinement is a practical engineering contribution, and the dual-identity real-time speaking/listening stack with physical deployment and perceptual studies is a meaningful step beyond monologue talking heads. Strengths include physical builds, explicit ablations of the design loop (Table I), comparison to manual design on a compact head, on-device FPS, and a reasonably sized user study. The fixed-topology premise and incomplete success rate limit how far the “wide range / scalable personalization” claim currently reaches, but the systems contribution remains substantial for robotics and HRI.

major comments (4)
  1. Abstract and §I claim automated synthesis for a “wide range of facial morphologies,” but Table I reports only 66.7% success (10/15), and App. D7 attributes the remaining failures to empty feasible sets under the fixed topology (insufficient internal clearance; multi-module interlocks that pose updates and γ-scheduling cannot resolve). Those are structural template failures, not mere optimizer misses. The central automation claim therefore needs either (i) a precise characterization of the morphology class for which the template is feasible, with failure rates broken down by geometry type, or (ii) an explicit limitation that topology search / reconfigurable hardware is required outside that class. Without this, “wide range” overstates what Alg. 1 delivers.
  2. §V-B, “Comparative efficiency against manual design”: the manual baseline is two senior designers on a single compact head (β=0.84), 22.8 h vs 11.7 min. This is useful as a case study but is too thin to support a general claim of superiority over manual mechanical design. At minimum, report designer protocols, success criteria, and preferably additional heads (including a failure-mode geometry from App. D7). Otherwise reframe the claim as a single-instance efficiency illustration rather than a comparative evaluation pillar.
  3. Table I: Global Joint-Opt has 0% SR and very large IV. That baseline is informative only if it is a fair formulation of joint optimization (soft collision penalties, same L-BFGS budget, same initialization). Please specify the full joint objective weights (App. D5), iteration budgets, and whether any multi-start or constraint-handling variants were tried. If the joint baseline is essentially unsolvable as posed, the hierarchical gain is partly architectural necessity rather than empirical dominance; that should be stated carefully so the 66.7% figure is not over-interpreted.
  4. §III-B and App. D6–D7: the 15-head evaluation set is central to the hardware claim, but the paper does not fully specify selection criteria, scale distribution, or how many of the eight visualized identities in Fig. A.5 are among the 10 successes. Please tabulate per-head outcomes (success/fail, final Exp, IV, outer-loop iterations) and relate failures to measurable geometric features (e.g., chin clearance, inter-module volume). Without that, reproducibility and external validity of the 66.7% rate remain limited.
minor comments (8)
  1. Eq. (1): L_amp = 1/α_k with a maximization intent via smaller L is clear, but the multi-trajectory case (K trajectories, shared or per-k Φ) should state whether Φ is shared across trajectories of the same module and how α_k are aggregated into the reported Exp = Σ α_k.
  2. Fig. 3 / Alg. 1: clarify whether eyes and jaw are ever re-optimized after coarse init when outer-loop collisions involve those modules, or whether only brow/mouth bases move.
  3. Table II: DualTalk is offline (uses partner future motion); mark this more prominently in the table caption so the RT column is not the only cue.
  4. Table IV total (25.84 h) assumes parallel Stages 2–3; state single-printer sequential time as well for readers estimating lab throughput.
  5. App. C Morpheus comparison (13.3% SR) is valuable; move a short summary into the main text near Table I so architectural sensitivity is not buried.
  6. Notation: p_base vs p^base, and α_limit vs α^limit, are used inconsistently between main text and Algorithm 1; unify.
  7. User study (Fig. 7): report statistical tests (e.g., paired tests or CI) for Ours vs ablations rather than means alone.
  8. Related work: briefly position against other parametric/retargetable facial robots beyond Morpheus to clarify novelty of the hierarchical collision loop.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: empirical systems paper with independent optimization objectives, collision metrics, and physical validation.

full rationale

The load-bearing claims (hierarchical kinematic synthesis under anatomy-guided volumes and AU trajectories, outer-loop MTV/QP collision refinement, 66.7% success on 15 heads vs. baselines in Table I, dual-identity audio-to-blendshape model, region-wise mapping) are defined by explicit objectives (Eq. 1, Eq. 2, Alg. 1) and measured by independent quantities (collision-free CAD assemblies, physical builds, LVE/MSE/PDD/SID/FDD, LSE-D/C, user ratings). Expressiveness coefficients α_k are optimized, not fitted then re-predicted; mapping MLPs are standard inverse calibration on robot-captured pairs, not circular proofs. Self-citations (Morpheus [111], DualTalk [83]) appear only as baselines or related work and are not used to force uniqueness or smuggle an ansatz into the central derivation. Failures (App. D7) are openly reported as empty feasible sets under the fixed topology, confirming the evaluation is falsifiable rather than definitional. The paper is self-contained against its own benchmarks.

Assumptions & free parameters 5 free parameters · 5 assumptions · 3 invented entities

The central automation claim rests on a fixed modular linkage template, anatomy-derived workspace bounds, AU-derived trajectory primitives, and a set of hand-chosen optimization weights and scheduling constants. None of these are free-form inventions of new physics; they are engineering priors. Free parameters are the optimizer weights and scheduling knobs; axioms are standard kinematics plus domain anatomy/FACS assumptions; invented entities are the paper’s template and hierarchical loop as design objects.

free parameters (5)
  • ω_amp, ω_fit, ω_man (objective weights)
    User-defined weights in Eq. 1 balancing amplitude, trajectory fit, and manipulability; tabulated as 1.0 / 10 / 0.1 in App. D4 without data-driven selection.
  • α_limit_init and γ (amplitude schedule)
    Initial amplitude cap 2.0 and decay factor 0.9 when pose update exceeds D_max; hand-chosen outer-loop controls that directly affect reported expressiveness and success.
  • D_max, ε (pose-update and clearance)
    Maximum accepted base-pose step 3.0 mm and safety clearance 0.5 mm in QP (Eq. 2); engineering thresholds that gate pose adjustment vs amplitude reduction.
  • β (global mesh scale) and template anthropometric offsets
    Per-head scale and template relative offsets initialize base poses; they encode the mean-face prior that later optimization only partially corrects.
  • Mapping MLP architecture and training hyperparameters
    Hidden sizes, dropout, learning rate, 3000 Latin-hypercube samples per head; fitted per physical robot and required for the physical-execution claim.
assumptions (5)
  • ad hoc to paper A fixed modular linkage topology (six-bar brows, four-bar lids/lips/jaw, five-bar corners, 2+1 neck) is an adequate universal template for diverse faces.
    Sec. III-A and App. D6–D7; topology is never searched. Failures when the constraint set is empty show this is load-bearing.
  • domain assumption Anatomy-guided feasible volumes (eyebrow planes, cheek slab, mid-sagittal lip corridor) bound bio-plausible motion.
    Sec. III-B2 and App. D2 citing Gray’s Anatomy and eyelid anthropometry [20,13].
  • domain assumption AU-derived canonical trajectories capture the semantic directions that mechanisms should maximize.
    Sec. III-B2 / App. D3; FACS [22] used as trajectory primitives rather than full muscle simulation.
  • standard math Spatial/planar linkage kinematics (Grübler–Kutzbach mobility, Newton–Raphson FK, closed-form four/five-bar solutions) correctly model the hardware.
    App. C full kinematic derivations used both for synthesis and motor IK.
  • domain assumption Dual-speaker audio alone (with turn gate) is sufficient to generate plausible listening and speaking facial motion without partner future motion.
    Sec. IV-A; contrasts with DualTalk’s offline partner-motion dependence.
invented entities (3)
  • Parametric linkage-driven mechanical face template
    purpose: Provide a single retargetable topology and actuator layout for automated scaling across faces.
    Introduced in Sec. III-A as the starting point of the pipeline; not independently validated outside this design family.
  • Hierarchical automatic design algorithm (inner kinematic opt + outer MTV/QP collision loop)
    purpose: Synthesize collision-free manufacturable mechanisms from 2D portraits.
    Algorithm 1 / Sec. III-B; empirical success is the only external handle.
  • Dual-identity conversational facial motion framework with turn-aware gate and region-wise motor mapping
    purpose: Produce real-time speaking and listening motion executable on the physical face.
    Sec. IV; evaluated on a custom dual-speaker blendshape dataset and user studies.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots." pith.science (2026). https://pith.science/paper/IJLQ7J33

@misc{pith2026260711688,
  author       = {Pith},
  title        = {Pith review of: Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IJLQ7J33}},
  note         = {Machine review of arXiv:2607.11688}
}
read the original abstract

Animatronic faces are a central component of socially interactive robots, enabling rich nonverbal communication through facial articulation. However, state-of-the-art animatronic faces are typically tailored systems: each new facial geometry requires extensive manual mechanical redesign, making large-scale personalization prohibitively slow and costly. In this work, we pursue automated and scalable mechanical face synthesis, aiming to rapidly generate a physically realizable facial mechanism for a wide range of facial geometries. We introduce a parametric, linkage-driven mechanical face template whose topology and actuator layout are explicitly parameterized to support systematic scaling and retargeting across diverse facial morphologies. Building on this template, we propose a hierarchical automatic design algorithm that takes a single 2D portrait as input, reconstructs a target 3D face, and synthesizes a collision-free, manufacturable internal mechanism. The algorithm combines anatomy-guided feasible motion volumes, Action Unit (AU)-derived trajectory-based expressiveness objectives, and a collision-driven outer-loop refinement strategy. Beyond hardware synthesis, we argue that future mechanical faces deployed at scale must engage in bidirectional, multi-turn conversation rather than functioning solely as speaking or listening heads. To this end, we develop a dual-identity conversational facial motion synthesis framework that jointly models speaking and listening behaviors from audio, producing temporally coherent 3D facial motion suitable for physical execution. We validate our system through extensive experiments, including (i) quantitative evaluation of automatic mechanism synthesis across diverse facial geometries, (ii) comparisons against manual mechanical design, (iii) benchmarks on conversational facial motion synthesis and real-time deployment, and (iv) perceptual user studies.

Figures

Figures reproduced from arXiv: 2607.11688 by the authors.

Figure 1
Figure 1. We demonstrate our end-to-end physical conversational face system across diverse multi-round interactions: (a) [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Mechanical face template. (a) Full assembly of the linkage-driven robotic face with a soft skin and a 3-DoF neck module. (b) Exploded view of the four modular facial mechanisms—eyebrow, eyes, mouth, and jaw. eyebrow ridge and the brow center. Both are realized by a spatial six-bar mechanism with 2 DoF. The eye module contains 8 DoF. It comprises four eyelid linkages (upper/lower lids for both eyes) and two rotationa… view at source ↗
Figure 3
Figure 3. Overview of the hierarchical automatic design pipeline. (a) Initialization: From a 2D portrait, we reconstruct a 3D head mesh, semantic landmarks, and initialize module base poses. (b) Inner loop: We perform module-wise kinematic synthesis under anatomy-guided feasible volumes and AU-derived trajectories to maximize expressiveness (illustrated with the mouth-corner five-bar mechanism). (c) Outer loop: We assemble th… view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Overview of interaction synthesis and control. (a) Given dual-speaker audio, the talking-head model predicts multi￾round facial motion for both speaking and listening. (b) The predicted coefficients are mapped to robot motor commands via region-wise regressors and neck…
Figure 5
Figure 5. Figure 5: (a) Algorithm design versatility: CAD Assemblies for Jack and Yoda. (b) Algorithm vs. Manual: Comparison of mouth-corner mechanisms (red circles) on a compact head. TABLE II: Comparison under speaker and listener modes on our dataset. “RT” means whether real-time use i…
Figure 6
Figure 6. Figure 6: End-to-end system visualization across interaction scenarios. (a) Human-Robot Interaction: Real-time conver￾sational gestures during a user-robot dialogue. The top row is the expression synthesis. (b) Dyadic Role-Play (Titanic): Reenacting the scene dynamics with expre…
Figure 7
Figure 7. Figure 7: User study results. (a) Overall interaction performance on four excerpts. (b) Mapping quality on three sentences. provide essential cues beyond the mouth region. We also evaluated mapping quality using three sentences [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

155 extracted references · 13 linked inside Pith

  1. [1]

    Designing of android head system by applying facial muscle mecha- nism of humans

    Ho Seok Ahn, Dong-Wook Lee, Dongwoon Choi, Duk- Yeon Lee, Manhong Hur, and Hogil Lee. Designing of android head system by applying facial muscle mecha- nism of humans. In2012 12th IEEE-RAS International Conference on Humanoid Robots (Humanoids 2012), pages 799–804. IEEE, 2012

  2. [2]

    The Design of an Expressive Humanlike Socially Assistive Robot.Journal of Mechanisms and Robotics, 1(1), 2008

    Brian Allison, Goldie Nejat, and Emmeline Kao. The Design of an Expressive Humanlike Socially Assistive Robot.Journal of Mechanisms and Robotics, 1(1), 2008

  3. [3]

    Linkedit: Interactive linkage editing using symbolic kinematics.ACM Transactions on Graphics (TOG), 34(4):1–8, 2015

    Moritz B ¨acher, Stelian Coros, and Bernhard Thomaszewski. Linkedit: Interactive linkage editing using symbolic kinematics.ACM Transactions on Graphics (TOG), 34(4):1–8, 2015

  4. [4]

    wav2vec 2.0: A framework for self- supervised learning of speech representations.Advances in neural information processing systems, 33:12449– 12460, 2020

    Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. wav2vec 2.0: A framework for self- supervised learning of speech representations.Advances in neural information processing systems, 33:12449– 12460, 2020

  5. [5]

    Control of facial expressions of the humanoid robot head ROMAN

    Karsten Berns and Jochen Hirth. Control of facial expressions of the humanoid robot head ROMAN. In 2006 IEEE/RSJ international conference on intelligent robots and systems, pages 3119–3124. IEEE, 2006

  6. [6]

    Physical face cloning.ACM Transac- tions on Graphics (TOG), 31(4):1–10, 2012

    Bernd Bickel, Peter Kaufmann, M ´elina Skouras, Bern- hard Thomaszewski, Derek Bradley, Thabo Beeler, Phil Jackson, Steve Marschner, Wojciech Matusik, and Markus Gross. Physical face cloning.ACM Transac- tions on Graphics (TOG), 31(4):1–10, 2012

  7. [7]

    The Cog project: Building a humanoid robot

    Rodney A Brooks, Cynthia Breazeal, Matthew Mar- janovi´c, Brian Scassellati, and Matthew M Williamson. The Cog project: Building a humanoid robot. In International workshop on computation for metaphors, analogy, and agents, pages 52–87. Springer, 1998

  8. [8]

    3D-printing of non-assembly, articulated mod- els.ACM Transactions on Graphics (TOG), 31(6):1–8, 2012

    Jacques Cal `ı, Dan A Calian, Cristina Amati, Re- becca Kleinberger, Anthony Steed, Jan Kautz, and Tim Weyrich. 3D-printing of non-assembly, articulated mod- els.ACM Transactions on Graphics (TOG), 31(6):1–8, 2012

Show all 155 references
  1. [9]

    Designing and fabricating mechanical automata from mocap sequences.ACM Transactions on Graphics (TOG), 32(6):1–11, 2013

    Duygu Ceylan, Wilmot Li, Niloy J Mitra, Maneesh Agrawala, and Mark Pauly. Designing and fabricating mechanical automata from mocap sequences.ACM Transactions on Graphics (TOG), 32(6):1–11, 2013

  2. [10]

    Smile like you mean it: Driving animatronic robotic face with learned models

    Boyuan Chen, Yuhang Hu, Lianfeng Li, Sara Cum- mings, and Hod Lipson. Smile like you mean it: Driving animatronic robotic face with learned models. In2021 IEEE International Conference on Robotics and Automation (ICRA), pages 2739–2746. IEEE, 2021

  3. [11]

    Cafe-talk: Generating 3d talking face animation with multimodal coarse-and fine- grained control.arXiv preprint arXiv:2503.14517, 2025

    Hejia Chen, Haoxian Zhang, Shoulong Zhang, Xiao- qiang Liu, Sisi Zhuang, Yuan Zhang, Pengfei Wan, Di Zhang, and Shuai Li. Cafe-talk: Generating 3d talking face animation with multimodal coarse-and fine- grained control.arXiv preprint arXiv:2503.14517, 2025

  4. [12]

    Towards Seamless Interaction: Causal Turn-Level Modeling of Interactive 3D Conversational Head Dynamics.arXiv preprint arXiv:2512.15340, 2025

    Junjie Chen, Fei Wang, Zhihao Huang, Qing Zhou, Kun Li, Dan Guo, Linfeng Zhang, and Xun Yang. Towards Seamless Interaction: Causal Turn-Level Modeling of Interactive 3D Conversational Head Dynamics.arXiv preprint arXiv:2512.15340, 2025

  5. [13]

    Three- dimensional anthropometric analysis of eyelid aging among Chinese women.Journal of Plastic, Reconstruc- tive & Aesthetic Surgery, 74(1):135–142, 2021

    Yuming Chong, Jingyu Li, Xinyu Liu, Xiaojun Wang, Jiuzuo Huang, Nanze Yu, and Xiao Long. Three- dimensional anthropometric analysis of eyelid aging among Chinese women.Journal of Plastic, Reconstruc- tive & Aesthetic Surgery, 74(1):135–142, 2021

  6. [14]

    ARTalk: Speech-Driven 3D Head Animation via Autoregressive Model.arXiv preprint arXiv:2502.20323, 2025

    Xuangeng Chu, Nabarun Goswami, Ziteng Cui, Hanqin Wang, and Tatsuya Harada. ARTalk: Speech-Driven 3D Head Animation via Autoregressive Model.arXiv preprint arXiv:2502.20323, 2025

  7. [15]

    UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speak- ing.arXiv preprint arXiv:2512.09327, 2025

    Xuangeng Chu, Ruicong Liu, Yifei Huang, Yun Liu, Yichen Peng, and Bo Zheng. UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speak- ing.arXiv preprint arXiv:2512.09327, 2025

  8. [16]

    Abel: integrating humanoid body, emotions, and time perception to investigate social interaction and human cognition.Applied Sciences, 11(3):1070, 2021

    Lorenzo Cominelli, Gustav Hoegen, and Danilo De Rossi. Abel: integrating humanoid body, emotions, and time perception to investigate social interaction and human cognition.Applied Sciences, 11(3):1070, 2021

  9. [17]

    Computa- tional design of mechanical characters.ACM Transac- tions on Graphics (TOG), 32(4):1–12, 2013

    Stelian Coros, Bernhard Thomaszewski, Gioacchino Noris, Shinjiro Sueda, Moira Forberg, Robert W Sum- ner, Wojciech Matusik, and Bernd Bickel. Computa- tional design of mechanical characters.ACM Transac- tions on Graphics (TOG), 32(4):1–12, 2013

  10. [18]

    Capture, learning, and synthesis of 3D speaking styles

    Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J Black. Capture, learning, and synthesis of 3D speaking styles. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10101–10111, 2019

  11. [19]

    Emo- tional speech-driven animation with content-emotion disentanglement

    Radek Dan ˇeˇcek, Kiran Chhatre, Shashank Tripathi, Yandong Wen, Michael Black, and Timo Bolkart. Emo- tional speech-driven animation with content-emotion disentanglement. InSIGGRAPH Asia 2023 Conference Papers, pages 1–13, 2023

  12. [20]

    Drake, A

    Richard L. Drake, A. Wayne V ogl, and Adam W. M. Mitchell.Gray’s Anatomy for Students. Churchill Livingstone, Philadelphia, PA, 3rd edition, 2015. ISBN 978-0702051319

  13. [21]

    Computational multicopter design

    Tao Du, Adriana Schulz, Bo Zhu, Bernd Bickel, and Wojciech Matusik. Computational multicopter design. ACM Transactions on Graphics (TOG), 35(6):1–10, 2016

  14. [22]

    Facial action coding system.Environmental Psychology & Nonverbal Behavior, 1978

    Paul Ekman and Wallace V Friesen. Facial action coding system.Environmental Psychology & Nonverbal Behavior, 1978

  15. [23]

    Unitalker: Scaling up audio-driven 3d facial animation through a unified model

    Xiangyu Fan, Jiaqi Li, Zhiqian Lin, Weiye Xiao, and Lei Yang. Unitalker: Scaling up audio-driven 3d facial animation through a unified model. InEuropean Con- ference on Computer Vision, pages 204–221. Springer, 2024

  16. [24]

    A Soft-Skin Facial Robot Capable of Real-Time Emotion-Driven Actuation Through Visual Perception

    Xuanhe Fan, Huijuan Zhao, Shuangjiang He, Li Li, and Li Yu. A Soft-Skin Facial Robot Capable of Real-Time Emotion-Driven Actuation Through Visual Perception. InInternational Conference on Intelligent Robotics and Applications, pages 54–66. Springer, 2025

  17. [25]

    Faceformer: Speech-driven 3d facial animation with transformers

    Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura. Faceformer: Speech-driven 3d facial animation with transformers. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 18770–18780, 2022

  18. [26]

    Facially expressive humanoid robotic face

    Zanwar Faraj, Mert Selamet, Carlos Morales, Patricio Torres, Maimuna Hossain, Boyuan Chen, and Hod Lipson. Facially expressive humanoid robotic face. HardwareX, 9:e00117, 2021

  19. [27]

    The rise of the roboid

    Leopoldina Fortunati, Alessandra Sorrentino, Laura Fiorini, and Filippo Cavallo. The rise of the roboid. International Journal of Social Robotics, 13(6):1457– 1471, 2021

  20. [28]

    Skaterbots: Optimization-based design and motion synthesis for robotic creatures with legs and wheels.ACM Trans- actions on Graphics (TOG), 37(4):1–12, 2018

    Moritz Geilinger, Roi Poranne, Ruta Desai, Bern- hard Thomaszewski, and Stelian Coros. Skaterbots: Optimization-based design and motion synthesis for robotic creatures with legs and wheels.ACM Trans- actions on Graphics (TOG), 37(4):1–12, 2018

  21. [29]

    Exploiting passive behaviours for diverse musical playing using the parametric hand

    Kieran Gilday, Dohyeon Pyeon, S Dhanush, Kyu-Jin Cho, and Josie Hughes. Exploiting passive behaviours for diverse musical playing using the parametric hand. Frontiers in Robotics and AI, 11:1463744, 2024

  22. [30]

    Embodied manipulation with past and future morphologies through an open parametric hand design

    Kieran Gilday, Chapa Sirithunge, Fumiya Iida, and Josie Hughes. Embodied manipulation with past and future morphologies through an open parametric hand design. Science Robotics, 10(102):eads6437, 2025

  23. [31]

    Erica: The erato intelligent conversational android

    Dylan F Glas, Takashi Minato, Carlos T Ishi, Tatsuya Kawahara, and Hiroshi Ishiguro. Erica: The erato intelligent conversational android. In2016 25th IEEE International symposium on robot and human inter- active communication (RO-MAN), pages 22–29. IEEE, 2016

  24. [32]

    Reinforcement learning for improving agent design.Artificial life, 25(4):352–365, 2019

    David Ha. Reinforcement learning for improving agent design.Artificial life, 25(4):352–365, 2019

  25. [33]

    Task-based limb optimization for legged robots

    Sehoon Ha, Stelian Coros, Alexander Alspach, Joohyung Kim, and Katsu Yamane. Task-based limb optimization for legged robots. In2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 2062–2068. IEEE, 2016

  26. [34]

    Learning human- like facial expressions for android Phillip K

    Ahsan Habib, Sumit K Das, Ioana-Corina Bogdan, David Hanson, and Dan O Popa. Learning human- like facial expressions for android Phillip K. Dick. In 2014 IEEE International Conference on Automation Science and Engineering (CASE), pages 1159–1165. IEEE, 2014

  27. [35]

    Identity Emulation (IE): Bio-inspired Facial Expression Interfaces for Emotive Robots

    David Hanson, Giovanni Pioggia, S Dinelli, Fabio Di Francesco, R Francesconi, and Danilo De Rossi. Identity Emulation (IE): Bio-inspired Facial Expression Interfaces for Emotive Robots. InAAAI Mobile Robot Competition, pages 72–82, 2002

  28. [36]

    Development of the face robot SAY A for rich facial expressions

    Takuya Hashimoto, Sachio Hiramatsu, and Hiroshi Kobayashi. Development of the face robot SAY A for rich facial expressions. In2006 International Confer- ence on Mechatronics and Automation, pages 25–30. IEEE, 2006

  29. [37]

    Dynamic display of facial expressions on the face robot made by using a life mask

    Takuya Hashimoto, Sachio Hiramatsu, and Hiroshi Kobayashi. Dynamic display of facial expressions on the face robot made by using a life mask. InHumanoids 2008-8th IEEE-RAS International Conference on Hu- manoid Robots, pages 521–526. IEEE, 2008

  30. [38]

    Morph: Design co-optimization with reinforcement learning via a dif- ferentiable hardware model proxy

    Zhanpeng He and Matei Ciocarlie. Morph: Design co-optimization with reinforcement learning via a dif- ferentiable hardware model proxy. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 7764–7771. IEEE, 2024

  31. [39]

    An iPhone Pro is All You Need: Mimicking Facial Expressions on an Android Robot Head

    Marcel Heisler and Christian Becker-Asano. An iPhone Pro is All You Need: Mimicking Facial Expressions on an Android Robot Head. In2025 20th ACM/IEEE International Conference on Human-Robot Interaction (HRI), pages 1347–1351. IEEE, 2025

  32. [40]

    Automatic design and manufacture of soft robots.IEEE Transactions on Robotics, 28(2):457–466, 2011

    Jonathan Hiller and Hod Lipson. Automatic design and manufacture of soft robots.IEEE Transactions on Robotics, 28(2):457–466, 2011

  33. [41]

    Human-robot facial coexpression.Science Robotics, 9 (88):eadi4724, 2024

    Yuhang Hu, Boyuan Chen, Jiong Lin, Yunzhe Wang, Yingke Wang, Cameron Mehlman, and Hod Lipson. Human-robot facial coexpression.Science Robotics, 9 (88):eadi4724, 2024

  34. [42]

    Learning realistic lip motions for humanoid face robots

    Yuhang Hu, Jiong Lin, Judah Allen Goldfeder, Philippe M Wyder, Yifeng Cao, Steven Tian, Yunzhe Wang, Jingran Wang, Mengmeng Wang, Jie Zeng, et al. Learning realistic lip motions for humanoid face robots. Science Robotics, 11(110):eadx3017, 2026

  35. [43]

    De- signing actuation systems for animatronic figures via globally optimal discrete search.ACM Transactions on Graphics (TOG), 40(4):1–10, 2021

    Simon Huber, Roi Poranne, and Stelian Coros. De- signing actuation systems for animatronic figures via globally optimal discrete search.ACM Transactions on Graphics (TOG), 40(4):1–10, 2021

  36. [44]

    Realistic child robot “affetto” for understanding the caregiver-child attachment relationship that guides the child development

    Hisashi Ishihara, Yuichiro Yoshikawa, and Minoru Asada. Realistic child robot “affetto” for understanding the caregiver-child attachment relationship that guides the child development. In2011 ieee international con- ference on development and learning (icdl), volume 2, pages 1...

  37. [45]

    Task-Based Design and Policy Co-Optimization for Tendon-driven Underactuated Kinematic Chains

    Sharfin Islam, Zhanpeng He, and Matei Ciocarlie. Task-Based Design and Policy Co-Optimization for Tendon-driven Underactuated Kinematic Chains. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 12016–12023. IEEE, 2024

  38. [46]

    Springer, 2006

    Kazuko Itoh, Hiroyasu Miwa, Massimiliano Zecca, Hideaki Takanobu, Stefano Roccella, Maria Chiara Car- rozza, Paolo Dario, and Atsuo Takanishi.Mechanical design of emotion expression humanoid robot we-4rii. Springer, 2006

  39. [47]

    Realtalk: Real-time and realistic audio-driven face generation with 3d facial prior-guided identity alignment network.arXiv preprint arXiv:2406.18284, 2024

    Xiaozhong Ji, Chuming Lin, Zhonggan Ding, Ying Tai, Junwei Zhu, Xiaobin Hu, Donghao Luo, Yanhao Ge, and Chengjie Wang. Realtalk: Real-time and realistic audio-driven face generation with 3d facial prior-guided identity alignment network.arXiv preprint arXiv:2406.18284, 2024

  40. [48]

    Cybernetic human HRP-4C

    Kenji Kaneko, Fumio Kanehiro, Mitsuharu Morisawa, Kanako Miura, Shin’ichiro Nakaoka, and Shuuji Kajita. Cybernetic human HRP-4C. In2009 9th IEEE-RAS International Conference on Humanoid Robots, pages 7–14. IEEE, 2009

  41. [49]

    Design Optimization of Wire Arrangement With Variable Re- lay Points in Numerical Simulation for Tendon-Driven Robots.IEEE Robotics and Automation Letters, 9(2): 1388–1395, 2023

    Kento Kawaharazuka, Shunnosuke Yoshimura, Temma Suzuki, Kei Okada, and Masayuki Inaba. Design Optimization of Wire Arrangement With Variable Re- lay Points in Numerical Simulation for Tendon-Driven Robots.IEEE Robotics and Automation Letters, 9(2): 1388–1395, 2023

  42. [50]

    Robot Design Optimization with Rotational and Prismatic Joints using Black-Box Multi-Objective Op- timization

    Kento Kawaharazuka, Kei Okada, and Masayuki In- aba. Robot Design Optimization with Rotational and Prismatic Joints using Black-Box Multi-Objective Op- timization. In2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 4571–

  43. [51]

    Study on face robot for active human interface-mechanisms of face robot and expression of 6 basic facial expressions

    Hiroshi Kobayashi and Fumio Hara. Study on face robot for active human interface-mechanisms of face robot and expression of 6 basic facial expressions. In Proceedings of 1993 2nd IEEE International Workshop on Robot and Human Communication, pages 276–281. IEEE, 1993

  44. [52]

    Study of a face robot platform as a kansei medium

    Hiroshi Kobayashi, T Tsuji, and K Kikuchi. Study of a face robot platform as a kansei medium. In2000 26th Annual Conference of the IEEE Industrial Electronics Society. IECON 2000. 2000 IEEE International Con- ference on Industrial Electronics, Control and Instru- mentation. 21...

  45. [53]

    Realization of realistic and rich facial expressions by face robot

    Hiroshi Kobayashi, Yoshiro Ichikawa, Masaru Senda, and Taichi Shiiba. Realization of realistic and rich facial expressions by face robot. InProceedings 2003 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2003)(Cat. No. 03CH37453), volume 2, pages 112...

  46. [54]

    Computational design of passive grippers.ACM Transactions on Graphics (TOG), 41(4): 2–12, 2022

    Milin Kodnongbua, Ian Good, Yu Lou, Jeffrey Lipton, and Adriana Schulz. Computational design of passive grippers.ACM Transactions on Graphics (TOG), 41(4): 2–12, 2022

  47. [55]

    Hunyuanvideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024

    Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024

  48. [56]

    Development of an android for emotional expression and human interaction

    Dong-Wook Lee, Tae-Geun Lee, B So, Moosung Choi, Eun-Cheol Shin, K Yang, Moon-Hong Baek, Hong- Seok Kim, and Ho-Gil Lee. Development of an android for emotional expression and human interaction. In Seventeenth world congress the international federation of automatic control, S...

  49. [57]

    J. P. Lewis, Ken ichi Anjyo, Taehyun Rhee, Mengjie Zhang, Fr´ed´eric H. Pighin, and Zhigang Deng. Practice and Theory of Blendshape Facial Models. InEuro- graphics, 2014

  50. [58]

    Driving Ani- matronic Robot Facial Expression From Speech.arXiv preprint arXiv:2403.12670, 2024

    Boren Li, Hang Li, and Hangxin Liu. Driving Ani- matronic Robot Facial Expression From Speech.arXiv preprint arXiv:2403.12670, 2024

  51. [59]

    Task-based design of cable-driven artic- ulated mechanisms

    Jian Li, Sheldon Andrews, Krisztian G Birkas, and Paul G Kry. Task-based design of cable-driven artic- ulated mechanisms. InProceedings of the 1st Annual ACM Symposium on Computational Fabrication, pages 1–12, 2017

  52. [60]

    X2C: A Dataset Featuring Nuanced Facial Expressions for Realistic Humanoid Imitation.arXiv preprint arXiv:2505.11146, 2025

    Peizhen Li, Longbing Cao, Xiao-Ming Wu, Runze Yang, and Xiaohan Yu. X2C: A Dataset Featuring Nuanced Facial Expressions for Realistic Humanoid Imitation.arXiv preprint arXiv:2505.11146, 2025

  53. [61]

    Learning a model of facial shape and expression from 4D scans.ACM Trans

    Tianye Li, Timo Bolkart, Michael J Black, Hao Li, and Javier Romero. Learning a model of facial shape and expression from 4D scans.ACM Trans. Graph., 36(6): 194–1, 2017

  54. [62]

    Modular design automation of the morphologies, controllers, and vision systems for intelligent robots: a survey.Visual Intelligence, 1(1):2, 2023

    Wenji Li, Zhaojun Wang, Ruitao Mai, Pengxiang Ren, Qinchang Zhang, Yutao Zhou, Ning Xu, JiaFan Zhuang, Bin Xin, Liang Gao, et al. Modular design automation of the morphologies, controllers, and vision systems for intelligent robots: a survey.Visual Intelligence, 1(1):2, 2023

  55. [63]

    Under- standing Embodied Reference with Touch-Line Trans- former

    Yang Li, Xiaoxue Chen, Hao Zhao, Jiangtao Gong, Guyue Zhou, Federico Rossano, and Yixin Zhu. Under- standing Embodied Reference with Touch-Line Trans- former. InICLR, 2023

  56. [64]

    The realization of robot theater: Humanoid robots and theatric perfor- mance

    Chyi-Yeu Lin, Chang-Kuo Tseng, Wei-Chung Teng, Wei-Chen Lee, Chung-Hsien Kuo, Hung-Yan Gu, Kuo- Liang Chung, and Chin-Shyurng Fahn. The realization of robot theater: Humanoid robots and theatric perfor- mance. In2009 International Conference on Advanced Robotics, pages 1–6. IEEE, 2009

  57. [65]

    Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models

    Gaojie Lin, Jianwen Jiang, Jiaqi Yang, Zerong Zheng, Chao Liang, Yuan Zhang, and Jingtuo Liu. Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models. InProceedings of the IEEE/CVF International Conference on Com- puter Vision, pages 13847–13858, 2025

  58. [66]

    Automatic design and manufacture of robotic lifeforms.Nature, 406 (6799):974–978, 2000

    Hod Lipson and Jordan B Pollack. Automatic design and manufacture of robotic lifeforms.Nature, 406 (6799):974–978, 2000

  59. [67]

    Emoface: Audio-driven emotional 3d face animation

    Chang Liu, Qunfen Lin, Zijiao Zeng, and Ye Pan. Emoface: Audio-driven emotional 3d face animation. In2024 IEEE Conference Virtual Reality and 3D User Interfaces (VR), pages 387–397. IEEE, 2024

  60. [68]

    Real-time robotic mirrored behavior of facial expressions and head motions based on lightweight networks.IEEE Internet of Things Journal, 10(2):1401– 1413, 2022

    Xiaofeng Liu, Yizhou Chen, Jie Li, and Angelo Can- gelosi. Real-time robotic mirrored behavior of facial expressions and head motions based on lightweight networks.IEEE Internet of Things Journal, 10(2):1401– 1413, 2022

  61. [69]

    Xiaofeng Liu, Rongrong Ni, Biao Yang, Siyang Song, and Angelo Cangelosi. Unlocking human-like facial expressions in humanoid robots: A novel approach for action unit driven facial expression disentangled synthesis.IEEE Transactions on Robotics, 40:3850– 3865, 2024

  62. [70]

    LSF- Animation: Label-Free Speech-Driven Facial Animation via Implicit Feature Representation.arXiv preprint arXiv:2510.21864, 2025

    Xin Lu, Chuanqing Zhuang, Chenxi Jin, Zhengda Lu, Yiqun Wang, Wu Liu, and Jun Xiao. LSF- Animation: Label-Free Speech-Driven Facial Animation via Implicit Feature Representation.arXiv preprint arXiv:2510.21864, 2025

  63. [71]

    MediaPipe: A Framework for Building Perception Pipelines.ArXiv, abs/1906.08172, 2019

    Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Guang Yong, Juhyun Lee, Wan-Teh Chang, Wei Hua, Manfred Georg, and Matthias Grundmann. MediaPipe: A Framework for Building Perception Pipelines.ArXiv, ab...

  64. [72]

    The bielefeld anthropomorphic robot head “Flobi”

    Ingo L ¨utkebohle, Frank Hegel, Simon Schulz, Matthias Hackel, Britta Wrede, Sven Wachsmuth, and Ger- hard Sagerer. The bielefeld anthropomorphic robot head “Flobi”. In2010 IEEE international conference on robotics and automation, pages 3384–3391. IEEE, 2010

  65. [73]

    Design and control co-optimization for automated design iteration of dexterous anthropomor- phic soft robotic hands

    Pragna Mannam, Xingyu Liu, Ding Zhao, Jean Oh, and Nancy Pollard. Design and control co-optimization for automated design iteration of dexterous anthropomor- phic soft robotic hands. In2024 IEEE 7th International Conference on Soft Robotics (RoboSoft), pages 332–

  66. [74]

    ChaCra: An Interactive Design System for Rapid Character Crafting

    Vittorio Megaro, Bernhard Thomaszewski, Damien Gauge, Eitan Grinspun, Stelian Coros, and Markus H Gross. ChaCra: An Interactive Design System for Rapid Character Crafting. InSymposium on Computer Animation, pages 123–130, 2014

  67. [75]

    Creating facial motions of Cybernetic Human HRP-4C

    Shin’ichiro Nakaoka, Fumio Kanehiro, Kanako Miura, Mitsuharu Morisawa, Kiyoshi Fujiwara, Kenji Kaneko, Shuuji Kajita, and Hirohisa Hirukawa. Creating facial motions of Cybernetic Human HRP-4C. In2009 9th IEEE-RAS International Conference on Humanoid Robots, pages 561–567. IEEE, 2009

  68. [76]

    From audio to photoreal embodiment: Synthe- sizing humans in conversations

    Evonne Ng, Javier Romero, Timur Bagautdinov, Shaojie Bai, Trevor Darrell, Angjoo Kanazawa, and Alexander Richard. From audio to photoreal embodiment: Synthe- sizing humans in conversations. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pa...

  69. [77]

    Celia Nieto Agraz, Pascal Hinrichs, Marco Eichelberg, and Andreas Hein. Is the robot spying on me? a study on perceived privacy in telepresence scenarios in a care setting with mobile and humanoid robots.International Journal of Social Robotics, 17(3):363–377, 2025

  70. [78]

    Design of android type humanoid robot Albert HUBO

    Jun-Ho Oh, David Hanson, Won-Sup Kim, Young Han, Jung-Yup Kim, and Ill-Woo Park. Design of android type humanoid robot Albert HUBO. In2006 IEEE/RSJ International Conference on Intelligent Robots and Sys- tems, pages 1428–1433. IEEE, 2006

  71. [79]

    FCL: A general purpose library for collision and proxim- ity queries

    Jia Pan, Sachin Chitta, and Dinesh Manocha. FCL: A general purpose library for collision and proxim- ity queries. In2012 IEEE international conference on robotics and automation, pages 3859–3866. IEEE, 2012

  72. [80]

    Self- talk: A self-supervised commutative training diagram to comprehend 3d talking faces

    Ziqiao Peng, Yihao Luo, Yue Shi, Hao Xu, Xiangyu Zhu, Hongyan Liu, Jun He, and Zhaoxin Fan. Self- talk: A self-supervised commutative training diagram to comprehend 3d talking faces. InProceedings of the 31st ACM International Conference on Multimedia, pages 5292–5301, 2023

  73. [81]

    Emotalk: Speech-driven emotional disentanglement for 3d face animation

    Ziqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu, Xi- angyu Zhu, Jun He, Hongyan Liu, and Zhaoxin Fan. Emotalk: Speech-driven emotional disentanglement for 3d face animation. InProceedings of the IEEE/CVF international conference on computer vision, pages 20687–20697, 2023

  74. [82]

    Synctalk: The devil is in the synchroniza- tion for talking head synthesis

    Ziqiao Peng, Wentao Hu, Yue Shi, Xiangyu Zhu, Xi- aomei Zhang, Hao Zhao, Jun He, Hongyan Liu, and Zhaoxin Fan. Synctalk: The devil is in the synchroniza- tion for talking head synthesis. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages...

  75. [83]

    Dualtalk: Dual- speaker interaction for 3d talking head conversations

    Ziqiao Peng, Yanbo Fan, Haoyu Wu, Xuan Wang, Hongyan Liu, Jun He, and Zhaoxin Fan. Dualtalk: Dual- speaker interaction for 3d talking head conversations. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 21055–21064, 2025

  76. [84]

    SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting.arXiv preprint arXiv:2506.14742, 2025

    Ziqiao Peng, Wentao Hu, Junyuan Ma, Xiangyu Zhu, Xiaomei Zhang, Hao Zhao, Hui Tian, Jun He, Hongyan Liu, and Zhaoxin Fan. SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting.arXiv preprint arXiv:2506.14742, 2025

  77. [85]

    Omnisync: Towards universal lip syn- chronization via diffusion transformers.arXiv preprint arXiv:2505.21448, 2025

    Ziqiao Peng, Jiwen Liu, Haoxian Zhang, Xiaoqiang Liu, Songlin Tang, Pengfei Wan, Di Zhang, Hongyan Liu, and Jun He. Omnisync: Towards universal lip syn- chronization via diffusion transformers.arXiv preprint arXiv:2505.21448, 2025

  78. [86]

    Analysis of suitable geometrical parameters for design- ing a tendon-driven under-actuated mechanical finger

    Francesco Penta, Cesare Rossi, and Sergio Savino. Analysis of suitable geometrical parameters for design- ing a tendon-driven under-actuated mechanical finger. Frontiers of Mechanical Engineering, 11(2):184–194, 2016

  79. [87]

    A lip sync expert is all you need for speech to lip generation in the wild

    KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Nam- boodiri, and CV Jawahar. A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM international conference on multimedia, pages 484–492, 2020

  80. [88]

    Meshtalk: 3d face animation from speech using cross-modality disentanglement

    Alexander Richard, Michael Zollh ¨ofer, Yandong Wen, Fernando De la Torre, and Yaser Sheikh. Meshtalk: 3d face animation from speech using cross-modality disentanglement. InProceedings of the IEEE/CVF in- ternational conference on computer vision, pages 1173– 1182, 2021

  81. [89]

    Jointly learning to construct and control agents using deep reinforcement learning

    Charles Schaff, David Yunis, Ayan Chakrabarti, and Matthew R Walter. Jointly learning to construct and control agents using deep reinforcement learning. In 2019 international conference on robotics and automa- tion (ICRA), pages 9798–9805. IEEE, 2019

  82. [90]

    A review on an AI-driven face robot for human-robot expression interaction.Science China Technological Sciences, 68(10):2010301, 2025

    Qincheng Sheng, Wei Tang, Hao Qin, Yujie Kong, Haokai Dai, Yiding Zhong, Yonghao Wang, Jun Zou, and Huayong Yang. A review on an AI-driven face robot for human-robot expression interaction.Science China Technological Sciences, 68(10):2010301, 2025

  83. [91]

    SAM Au- dio: Segment Anything in Audio.arXiv preprint arXiv:2512.18099, 2025

    Bowen Shi, Andros Tjandra, John Hoffman, Helin Wang, Yi-Chiao Wu, Luya Gao, Julius Richter, Matt Le, Apoorv Vyas, Sanyuan Chen, et al. SAM Au- dio: Segment Anything in Audio.arXiv preprint arXiv:2512.18099, 2025

  84. [92]

    Transnet v2: An effective deep network architecture for fast shot tran- sition detection

    Tom ´as Soucek and Jakub Lokoc. Transnet v2: An effective deep network architecture for fast shot tran- sition detection. InProceedings of the 32nd ACM International Conference on Multimedia, pages 11218– 11221, 2024

  85. [93]

    Knowledge- driven automated design of industrial robots: A unified graph-based framework with multi-engine reasoning

    Tao Sun, Bo Wang, and Xinming Huo. Knowledge- driven automated design of industrial robots: A unified graph-based framework with multi-engine reasoning. Advanced Engineering Informatics, 69:103995, 2026

  86. [94]

    Computational design of linkage-based charac- ters.ACM Transactions on Graphics (TOG), 33(4):1–9, 2014

    Bernhard Thomaszewski, Stelian Coros, Damien Gauge, Vittorio Megaro, Eitan Grinspun, and Markus Gross. Computational design of linkage-based charac- ters.ACM Transactions on Graphics (TOG), 33(4):1–9, 2014

  87. [95]

    Dim: Dyadic interaction modeling for social behavior generation

    Minh Tran, Di Chang, Maksim Siniukov, and Moham- mad Soleymani. Dim: Dyadic interaction modeling for social behavior generation. InEuropean Conference on Computer Vision, pages 484–503. Springer, 2024

  88. [96]

    Neural graph evolution: Towards efficient automatic robot design.arXiv preprint arXiv:1906.05370, 2019

    Tingwu Wang, Yuhao Zhou, Sanja Fidler, and Jimmy Ba. Neural graph evolution: Towards efficient automatic robot design.arXiv preprint arXiv:1906.05370, 2019

  89. [97]

    Development of the humanoid head portrait robot system with flexible face and expression

    Wu Weiguo, Men Qingmei, and Wang Yu. Development of the humanoid head portrait robot system with flexible face and expression. In2004 IEEE International Con- ference on Robotics and Biomimetics, pages 757–762. IEEE, 2004

  90. [98]

    Retargeting human facial expres- sion to human-like robotic face through neural network surrogate-based optimization

    Bowen Wu, Chaoran Liu, Carlos T Ishi, Takashi Minato, and Hiroshi Ishiguro. Retargeting human facial expres- sion to human-like robotic face through neural network surrogate-based optimization. In2024 IEEE/RSJ Inter- national Conference on Intelligent Robots and Systems (IROS)...

  91. [99]

    Codetalker: Speech-driven 3d facial animation with discrete motion prior

    Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang, and Tien-Tsin Wong. Codetalker: Speech-driven 3d facial animation with discrete motion prior. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12780–12790, 2023

  92. [100]

    Learning to fly: computational controller design for hybrid uavs with reinforcement learning.ACM Transactions on Graphics (TOG), 38(4):1–12, 2019

    Jie Xu, Tao Du, Michael Foshey, Beichen Li, Bo Zhu, Adriana Schulz, and Wojciech Matusik. Learning to fly: computational controller design for hybrid uavs with reinforcement learning.ACM Transactions on Graphics (TOG), 38(4):1–12, 2019

  93. [101]

    An End-to-End Differentiable Framework for Contact- Aware Robot Design

    Jie Xu, Tao Chen, Lara Zlokapa, Michael Foshey, Wojciech Matusik, Shinjiro Sueda, and Pulkit Agrawal. An End-to-End Differentiable Framework for Contact- Aware Robot Design. InRobotics: Science & Systems, 2021

  94. [102]

    SingingBot: An Avatar-Driven System for Robotic Face Singing Per- formance.arXiv preprint arXiv:2601.02125, 2026

    Zhuoxiong Xu, Xuanchen Li, Yuhao Cheng, Fei Xu, Yichao Yan, and Xiaokang Yang. SingingBot: An Avatar-Driven System for Robotic Face Singing Per- formance.arXiv preprint arXiv:2601.02125, 2026

  95. [103]

    Optimiz- ing facial expressions of an android robot effectively: a bayesian optimization approach

    Dongsheng Yang, Wataru Sato, Qianying Liu, Takashi Minato, Shushi Namba, and Shin’ya Nishida. Optimiz- ing facial expressions of an android robot effectively: a bayesian optimization approach. In2022 IEEE-RAS 21st International Conference on Humanoid Robots (Humanoids), pages ...

  96. [104]

    HAPI: A Model for Learning Robot Facial Expressions from Human Preferences.arXiv preprint arXiv:2503.17046, 2025

    Dongsheng Yang, Qianying Liu, Wataru Sato, Takashi Minato, Chaoran Liu, and Shin’ya Nishida. HAPI: A Model for Learning Robot Facial Expressions from Human Preferences.arXiv preprint arXiv:2503.17046, 2025

  97. [105]

    Co-Design of Soft Gripper with Neural Physics.arXiv preprint arXiv:2505.20404, 2025

    Sha Yi, Xueqian Bai, Adabhav Singh, Jianglong Ye, Michael T Tolley, and Xiaolong Wang. Co-Design of Soft Gripper with Neural Physics.arXiv preprint arXiv:2505.20404, 2025

  98. [106]

    ExFace: Expressive Facial Control for Humanoid Robots with Diffusion Transformers and Bootstrap Training.arXiv preprint arXiv:2504.14477, 2025

    Dong Zhang, Jingwei Peng, Yuyang Jiao, Jiayuan Gu, Jingyi Yu, and Jiahao Chen. ExFace: Expressive Facial Control for Humanoid Robots with Diffusion Transformers and Bootstrap Training.arXiv preprint arXiv:2504.14477, 2025

  99. [107]

    Functionality-aware retargeting of mechanisms to 3D shapes.ACM Transactions on Graphics (TOG), 36(4):1–13, 2017

    Ran Zhang, Thomas Auzinger, Duygu Ceylan, Wilmot Li, and Bernd Bickel. Functionality-aware retargeting of mechanisms to 3D shapes.ACM Transactions on Graphics (TOG), 36(4):1–13, 2017

  100. [108]

    Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face anima- tion

    Wenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang, Xi Shen, Yu Guo, Ying Shan, and Fei Wang. Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face anima- tion. InProceedings of the IEEE/CVF conference on computer vision and p...

  101. [109]

    Fabg: End-to-end imitation learning for embodied affective human-robot interaction.arXiv preprint arXiv:2503.01363, 2025

    Yanghai Zhang, Changyi Liu, Keting Fu, Wenbin Zhou, Qingdu Li, and Jianwei Zhang. Fabg: End-to-end imitation learning for embodied affective human-robot interaction.arXiv preprint arXiv:2503.01363, 2025

  102. [110]

    Shoulder- Shot: Generating Over-the-Shoulder Dialogue Videos

    Yuang Zhang, Junqi Cheng, Haoyu Zhao, Jiaxi Gu, Fangyuan Zou, Zenghui Lu, and Peng Shu. Shoulder- Shot: Generating Over-the-Shoulder Dialogue Videos. arXiv preprint arXiv:2508.07597, 2025

  103. [111]

    Morpheus: A Neural-driven Animatronic Face with Hybrid Actuation and Diverse Emotion Con- trol.arXiv preprint arXiv:2507.16645, 2025

    Zongzheng Zhang, Jiawen Yang, Ziqiao Peng, Meng Yang, Jianzhu Ma, Lin Cheng, Huazhe Xu, Hang Zhao, and Hao Zhao. Morpheus: A Neural-driven Animatronic Face with Hybrid Actuation and Diverse Emotion Con- trol.arXiv preprint arXiv:2507.16645, 2025

  104. [112]

    Robogrammar: graph grammar for terrain-optimized robot design.ACM Transactions on Graphics (TOG), 39(6):1–16, 2020

    Allan Zhao, Jie Xu, Mina Konakovi ´c-Lukovi´c, Josephine Hughes, Andrew Spielberg, Daniela Rus, and Wojciech Matusik. Robogrammar: graph grammar for terrain-optimized robot design.ACM Transactions on Graphics (TOG), 39(6):1–16, 2020

  105. [113]

    Learning to draw sight lines.International Journal of Computer Vision, 128(5):1076–1100, 2020

    Hao Zhao, Ming Lu, Anbang Yao, Yurong Chen, and Li Zhang. Learning to draw sight lines.International Journal of Computer Vision, 128(5):1076–1100, 2020

  106. [114]

    Interactive conversational head generation

    Mohan Zhou, Yalong Bai, Wei Zhang, Ting Yao, and Tiejun Zhao. Interactive conversational head generation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

  107. [115]

    Motion-guided mechanical toy modeling.ACM Transactions on Graphics (TOG), 31(6):1–10, 2012

    Lifeng Zhu, Weiwei Xu, John Snyder, Yang Liu, Guop- ing Wang, and Baining Guo. Motion-guided mechanical toy modeling.ACM Transactions on Graphics (TOG), 31(6):1–10, 2012

  108. [116]

    INFP: Audio- driven interactive head generation in dyadic conver- sations

    Yongming Zhu, Longhao Zhang, Zhengkun Rong, Tian- shu Hu, Shuang Liang, and Zhipeng Ge. INFP: Audio- driven interactive head generation in dyadic conver- sations. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 10667–10677, 2025

  109. [117]

    Awakening Facial Emotional Expressions in Human- Robot.arXiv preprint arXiv:2510.23059, 2025

    Yongtong Zhu, Lei Li, Iggy Qian, WenBin Zhou, Ye Yuan, Qingdu Li, Na Liu, and Jianwei Zhang. Awakening Facial Emotional Expressions in Human- Robot.arXiv preprint arXiv:2510.23059, 2025

  110. [118]

    Towards Metrical Reconstruction of Human Faces.Eu- ropean Conference on Computer Vision, 2022

    Wojciech Zielonka, Timo Bolkart, and Justus Thies. Towards Metrical Reconstruction of Human Faces.Eu- ropean Conference on Computer Vision, 2022. APPENDIX A. Overview This appendix provides supplementary technical details, mathematical derivations, and extended experimental re...

  111. [119]

    Mechanical Face Platform:Robotic face design has evolved significantly to enhance human-robot interaction, primarily driven by advancements in actuation modalities. Early research utilized flexible micro-actuators and pneumatic artificial muscles to replicate FACS-based motion...

  112. [120]

    Complementary tools support rapid fabri- cation [74, 8] and the retargeting of mechanism templates to arbitrary 3D geometries [107]

    Automatic Robot Design:Computer graphics research has extensively explored synthesizing mechanisms from mo- tion specifications, ranging from planar linkages [17, 3] and linkage-based characters [94] to automata derived from motion capture [9, 115]. Complementary tools support...

  113. [121]

    Consequently, 3D talking head synthe- sis [18, 88, 14] is preferred for robotic actuation [111]

    Talking Head Generation and Expression Mapping: Although 2D audio-driven talking face generation achieves impressive visual fidelity [87, 108, 55, 85, 65], these methods typically suffer from high generation latency [114, 116, 102, 42] and fail to provide robust mapping perspe...

  114. [122]

    A.2(a) shows the eyebrow mech- anism, implemented as a spatial six-bar linkage with left/right symmetry

    Eyebrow module:Fig. A.2(a) shows the eyebrow mech- anism, implemented as a spatial six-bar linkage with left/right symmetry. From the ground to the output member, the closed chain can be viewed as a kinematic sequenceR–S–S–S–R– R, whereRdenotes a revolute joint andSdenotes a s...

  115. [123]

    Eyes module:As shown in Fig. A.2(b), each eye inte- grates four four-bar linkages: two four-bar mechanisms drive the upper and lower eyelids for blinking, and two additional four-bar mechanisms actuate the eyeball for left–right (yaw) motion. Altogether, the binocular eye modu...

  116. [124]

    elbow- up

    Mouth module:The mouth module comprises four sub- components: the upper lip, lower lip, left mouth corner, and right mouth corner (Fig. A.2(c)). The upper and lower lips are each actuated by a symmetric four-bar linkage; the output attachment pointEcouples to the skin to gener...

  117. [125]

    A.2(d), jaw open- ing/closing is actuated by a compact four-bar linkage

    Jaw module:As illustrated in Fig. A.2(d), jaw open- ing/closing is actuated by a compact four-bar linkage. Points AandDare fixed ground pivots, while the remaining links transmit the actuator motion to the jaw output

  118. [126]

    Neck module:This section derives the closed-formin- verse kinematics(IK) for the neck mechanism in Fig. A.3. The mechanism realizes a three-axis rotational motion of the moving platform. Given a desired platform orientationR, we compute the global positions of the two platform...

  119. [127]

    Morpheus decomposes facial actuation into multiple modules

    Architecture matters:To examine how mechanism ar- chitecture affects scalability, we additionally apply the same automatic design algorithm to an open-source robotic head, Morpheus [111]. Morpheus decomposes facial actuation into multiple modules. The eyebrow module separates ...

  120. [128]

    Initialization of geometric parameters and layout:To ensure a physically plausible starting point for the hierarchical optimization, we analytically derive the initial spatial pose and geometric parameters of each module directly from the reconstructed 3D semantic landmarks. E...

  121. [129]

    Determine anatomy-guided feasible motion volumes: We instantiate the anatomy-guided feasible motion volumes only for the actively optimized regions (eyebrow and mouth), providing explicit geometric bounds used to constructΩ feas (Fig. A.4(b)). Eyebrow volume.Following the anth...

  122. [130]

    Specifically, we derive a set of canonical motion trajectories from AU-based facial motion analysis and use them as kinematic targets for synthesis (Fig

    Determine AU-derived trajectories:In the inner-loop optimization, we impose trajectory constraints only on the actively optimized facial modules (eyebrow, lips, and mouth corners), since the eye and jaw modules are fixed after coarse initialization. Specifically, we derive a s...

  123. [131]

    A.1, including settings for the inner loop, the outer-loop collision-driven refinement, and the L-BFGS solver

    Optimization parameters:We summarize the key hy- perparameters used in our hierarchical automatic design opti- mization in Tab. A.1, including settings for the inner loop, the outer-loop collision-driven refinement, and the L-BFGS solver

  124. [132]

    This part details the mathematical formulation and algorithmic procedures for each baseline

    Baselines:To rigorously evaluate the efficacy of our pro- posed hierarchical optimization framework, we implemented three comparative baselines. This part details the mathematical formulation and algorithmic procedures for each baseline. Local-only optimization (Ours w/o outer...

  125. [133]

    Diverse facial mechanism synthesis visualization:To strictly validate the geometric versatility of our hierarchical optimization framework, we conducted a batch synthesis ex- periment on a diverse dataset of eight 3D portraits (Fig. A.5). This set encompasses both photorealist...

  126. [134]

    Template-infeasible morphology.In some extreme ge- ometries, the design becomesstructurally infeasibleunder our current mechanism template

    Failure cases:Although our hierarchical pipeline is robust across a wide range of face geometries, it can fail when the target morphology violates the feasibility assumptions of the fixed mechanism template. Template-infeasible morphology.In some extreme ge- ometries, the desi...

  127. [135]

    3D printed materials:All rigid mechanical components and the supporting shell are fabricated using Polylactic Acid (PLA) on a commercial high-speed FDM printer (Bambu Lab P1S). To ensure precise assembly and low-friction joint articulation without extensive post-processing, we...

  128. [136]

    We utilize a standard RTV-2 molding silicone

    Silicone face skin:The facial skin is fabricated using a double-mold casting technique. We utilize a standard RTV-2 molding silicone. The fabrication process proceeds as follows: Mold Preparation:A rigid outer mold (defining the facial surface) and an inner core (defining the ...

  129. [137]

    The neck mechanism, which requires a higher holding torque, is driven by three NEMA 17 stepper motors

    Actuator choice and noise testing:To balance torque density with the spatial constraints of the compact cephalic volume, we select GuoHua 9g digital micro-servos (180 ◦ range). The neck mechanism, which requires a higher holding torque, is driven by three NEMA 17 stepper motor...

  130. [138]

    Network architecture:In this section, we elucidate the comprehensive architectural specifications of our Unified Dual-Speaker Interaction Framework (Fig. 4(a)). Dual-speaker joint encoder.The conversational motion synthesis framework is underpinned by a Dual-Speaker Joint Enco...

  131. [139]

    Transformer backbone.The core sequence modeling is performed by a Transformer architecture comprising an En- coder and a Decoder, both configured withN= 3layers

    and shift attention to the interlocutor’s audio for reactive listening behaviors when passive (g A ≈0). Transformer backbone.The core sequence modeling is performed by a Transformer architecture comprising an En- coder and a Decoder, both configured withN= 3layers. •Encoder:Th...

  132. [140]

    Each frame is parameterized by a 55-D vector yt = [b t;r t]∈R 55, whereb t ∈R 52 are blendshape coefficients, andr t ∈R 3 denotes head rotation

    Loss function:We train the network to predict a tem- porally coherent sequence of motion parameters from dual- speaker audio. Each frame is parameterized by a 55-D vector yt = [b t;r t]∈R 55, whereb t ∈R 52 are blendshape coefficients, andr t ∈R 3 denotes head rotation. Our ob...

  133. [141]

    The training curriculum spans a total of 200 epochs, processing data in mini-batches of size 32

    Training details:The network parameters are optimized using the Adam algorithm, initialized with a constant learning rate of1×10 −4. The training curriculum spans a total of 200 epochs, processing data in mini-batches of size 32. We employ a distributed training strategy on a ...

  134. [142]

    Lip Vertex Error (LVE,↓).Measures lip-synchronization accuracy by calculating the geometric deviation of the mouth region

    Evaluation metrics:We provide detailed definitions of the quantitative metrics employed to assess the performance of both speaker-centric synchronization and listener-centric reactive behaviors. Lip Vertex Error (LVE,↓).Measures lip-synchronization accuracy by calculating the ...

  135. [143]

    Dataset details:We build a dual-speaker interaction dataset from publicly available YouTube interview-style videos and RealTalk [47], which provide diverse face-to-face conver- sational behaviors. We select videos where (i) both speak- ers are clearly visible for most of the c...

  136. [144]

    Landmark representation.We extractD=4682D facial landmarks using MediaPipe [71]

    Baseline details:We provide implementation details for the baselines. Landmark representation.We extractD=4682D facial landmarks using MediaPipe [71]. We denote the landmark matrix asL t ={(x t,j, yt,j)}D j=1 ∈R D×2, and its vectorized form asl t = vec(Lt)∈R 2D. We usel H t fo...

  137. [145]

    For a given utterance, we extract short overlapping temporal windows from the audio stream and the corresponding robot’s mouth-region video frames

    Evaluation metrics:We evaluate audio–visual lip syn- chronization using Lip Sync Error Distance (LSE-D) and Lip Sync Error Confidence (LSE-C), which are computed with a pre-trained audio–visual synchronization network following the protocol used in prior work [87]. For a given...

  138. [146]

    Semantic region-wise mapping network:We implement a Semantic Region-wise Mapping Network based on a Multi- Layer Perceptron (MLP) to regress motor control signals from facial blendshape coefficients. We specifically detail the mapping architecture for the mouth and jaw regions...

  139. [147]

    In this scenario, a human operator engaged the robotic head in a semi-structured philosophical dialogue

    Real-time human-robot interaction:To strictly validate the system’s cross-lingual generalization capabilities and real- time inference latency, we conduct a live interaction ex- periment. In this scenario, a human operator engaged the robotic head in a semi-structured philosop...

  140. [148]

    In these experiments, two robotic heads driven by distinct personality profiles perform classic cinematic narratives

    Dyadic robot-robot dialogue:To demonstrate the frame- work’s capacity for autonomous dual-agent synchronization, we orchestrate two ”Physical Reenactment” scenarios. In these experiments, two robotic heads driven by distinct personality profiles perform classic cinematic narra...

  141. [149]

    In this setup, a Digital Avatar (displayed on a screen) interacts with the Physical Robot (Fig

    Hybrid avatar-robot interaction:To evaluate the sys- tem’s versatility across heterogeneous embodiments, we con- struct a mixed-reality scenario. In this setup, a Digital Avatar (displayed on a screen) interacts with the Physical Robot (Fig. A.6 (b)). Transcript: The Choice to...

  142. [150]

    For each excerpt, participants were first shown the reference movie clip and then evaluated five robot interaction variants presented in a randomized order

    Study 1: Interaction performance on movie dialogue excerpts:To evaluate the perceived quality of our real-time interaction pipeline, we conducted a perceptual study using four representative dialogue excerpts drawn fromTitanic,Star Wars,The Lord of the Rings, and a Chinese fol...

  143. [151]

    Kids are sitting by the door

    Study 2: Perceived mapping quality.:We conducted a second perceptual study to isolate the effect of differ- ent mapping strategies on full-face expressiveness. We used three short sentences:“Kids are sitting by the door”,“I am going to the store”, and“I lost my keys”. For each...

  144. [152]

    While effective, this approach places a substantial burden on the algorithmic synthesis to resolve both global spatial alignment and local kinematic performance simultaneously

    Mechanical reconfigurability and modular extensibility: Currently, our hierarchical framework primarily relies com- putational optimization to adapt the mechanism layout to diverse facial topologies. While effective, this approach places a substantial burden on the algorithmic...

  145. [153]

    Unified end-to-end learning framework:Our system op- erates as a modular pipeline where distinct components—such as audio feature extraction, expression coefficient prediction, and motion retargeting—are serialized. A fundamental lim- itation of this cascaded architecture is t...

  146. [154]

    Digital Twin

    Differentiable simulation and sim-to-real adaptation: A significant bottleneck in current robotic facial expression generation is the scarcity of high-quality, paired real-world data. Furthermore, the physical coupling between rigid me- chanical structures and hyper-elastic si...

  147. [155]

    However, human communication is inherently holistic, relying heavily on the coordination of facial cues with body posture and manual gestures to convey intent and emotion

    Towards full-body embodiment and interaction:Our current implementation focuses exclusively on facial expres- sions and neck movements. However, human communication is inherently holistic, relying heavily on the coordination of facial cues with body posture and manual gestures...

Pith tools

Reviewed July 14, 2026 · model on record in the stance chip above.