Pith. sign in

REVIEW 9 cited by

Body of Her: A Preliminary Study on End-to-End Humanoid Agent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.02879 v1 pith:VGL5Q6ZA submitted 2024-08-06 cs.CV

Body of Her: A Preliminary Study on End-to-End Humanoid Agent

classification cs.CV
keywords agenthumanoidmodelend-to-endmanipulationaudiobodycapable
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Interactive virtual humanoid agent is a crucial interface with the physical world. A relatively complete humanoid agent first needs to have face and body, then possess both verbal and non-verbal (such as eye contact, facial expression, lip motion, gesture, and manipulation) abilities, and finally, it is capable of real-time duplex communication, e.g., the ability to actively interrupt conversations. Most prior systems typically only consider a subset of these elements, leaving a gap from realistic humanoid agent. In this work, we propose a real-time, duplex, interactive end-to-end network capable of modeling realistic agent behaviors, including speech, full-body movements for talking, responding, idling, and manipulation. This system is a multimodal model integrating audio and visual inputs, extended from a pre-trained large language model (LLM). We collect approximately 200,000 hours of audio, around 130,000 hours of video data, and about 20,000 alignment samples to build the model. The final model demonstrates capabilities that are difficult to achieve in previous systems, such as generalized object manipulation. This work performs a preliminary exploration of the end-to-end approach in this field, aiming to inspire further research towards scaling up.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

    cs.CV 2026-06 unverdicted novelty 6.0

    Wan-Streamer is a unified end-to-end Transformer for low-latency streaming audio-visual interaction using block-causal attention on interleaved multimodal tokens.

  2. Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length

    cs.CV 2025-12 conditional novelty 6.0

    Live Avatar enables 45 FPS real-time streaming infinite-length audio-driven avatar generation from a 14B diffusion model via distillation and timestep-forcing pipeline parallelism.

  3. Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length

    cs.CV 2025-12 conditional novelty 6.0

    Live Avatar reports real-time streamable generation from a 14B audio-driven diffusion model at ~20 FPS on 5 H800s with stable identity over 10,000 seconds.

  4. OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation

    cs.CV 2025-10 conditional novelty 6.0

    A single autoregressive diffusion model, trained on a new 286-hour SMPL-X dataset, generates whole-body motion from text, audio, and spatial-temporal control signals, with reference-motion conditioning.

  5. Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

    cs.CV 2026-06 unverdicted novelty 5.0

    Wan-Streamer is a unified Transformer model for low-latency streaming audio-visual interaction that jointly handles perception, reasoning, generation, and timing without external modules.

  6. Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

    cs.CV 2026-06 unverdicted novelty 5.0

    Wan-Streamer presents a unified end-to-end Transformer for low-latency multimodal streaming interaction without external modules.

  7. MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

    cs.CV 2025-08 reject novelty 5.0

    A new autoregressive video-generation framework for interactive digital humans, with a 64x compression autoencoder and a diffusion renderer, claims real-time multimodal control but is only demonstrated for audio input.

  8. Wan-Streamer v0.2: Higher Resolution, Same Latency

    cs.CV 2026-07 conditional novelty 4.0

    Wan-Streamer v0.2 upgrades native-streaming audio-visual interaction to 640×368 at 25 FPS with unchanged ~200 ms model-side latency via a single-GPU thinker and multi-GPU Ulysses-style performer.

  9. Video = World + Event Stream

    cs.CV 2026-07 reject novelty 3.0

    Wan-Streamer v0.3 reformulates streaming video as a persistent world plus a changing event stream and extends an agent's output with open-vocabulary behavior, without reporting any quantitative results.