Pith. sign in

REVIEW 24 cited by

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.15111 v1 pith:M362KP5L submitted 2025-01-25 cs.CV

classification cs.CV
keywords human-centricscenesunderstandinghumanomnimodelbranchesfeaturesindividuals
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In human-centric scenes, the ability to simultaneously understand visual and auditory information is crucial. While recent omni models can process multiple modalities, they generally lack effectiveness in human-centric scenes due to the absence of large-scale, specialized datasets and non-targeted architectures. In this work, we developed HumanOmni, the industry's first human-centric Omni-multimodal large language model. We constructed a dataset containing over 2.4 million human-centric video clips with detailed captions and more than 14 million instructions, facilitating the understanding of diverse human-centric scenes. HumanOmni includes three specialized branches for understanding different types of scenes. It adaptively fuses features from these branches based on user instructions, significantly enhancing visual understanding in scenes centered around individuals. Moreover, HumanOmni integrates audio features to ensure a comprehensive understanding of environments and individuals. Our experiments validate HumanOmni's advanced capabilities in handling human-centric scenes across a variety of tasks, including emotion recognition, facial expression description, and action understanding. Our model will be open-sourced to facilitate further development and collaboration within both academia and industry.

Discussion (0). Sign in to comment.

Forward citations

Cited by 24 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos

    cs.CV 2026-05 unverdicted novelty 8.0 of 10

    TraceAV-Bench is the first benchmark for multi-hop trajectory reasoning over long audio-visual videos, showing top models reach only 51-68% accuracy with substantial room for improvement.

  2. InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring

    cs.AI 2026-07 conditional novelty 7.0 of 10

    A new in-cabin dataset pairs RGB/IR video, audio, and Chinese dialogue text with emotion, fatigue, and distraction labels, and baselines show fusion beats single modalities on the Chinese partition.

  3. AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    AVI-Bench is a cognitively inspired benchmark that evaluates Omni-MLLMs on joint audio-visual tasks and reveals substantial limitations in current models.

  4. GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    GRASP is a large-scale dataset and benchmark for social reasoning grounded in gaze and gesture events in multi-person videos, with Social Grounding Reward (SGR) proposed to improve model performance on GRASP-Bench.

  5. Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions

    cs.CV 2026-04 conditional novelty 7.0 of 10

    Creates the first egocentric screen-view movie emotion benchmark and demonstrates that cinematic models drop sharply in Macro-F1 on realistic robot-like viewing conditions while domain-specific training improves robustness.

  6. SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    SVAgent improves long video question answering by constructing storylines via multi-agent collaboration and aligning cross-modal predictions for more robust, human-like reasoning.

  7. Omni-Perception Policy Optimization for Multimodal Emotion Reasoning

    cs.AI 2026-06 conditional novelty 6.0 of 10

    A cue-coverage reward plus a modality-token KL penalty makes multimodal emotion-reasoning models cite more real visual/audio evidence and hallucinate less, with reported SoTA on emotion benchmarks.

  8. Omni-Perception Policy Optimization for Multimodal Emotion Reasoning

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    OPPO applies RL with an Omni-Perception Reward and masked-input KL loss to boost cue utilization and suppress hallucinations in emotion reasoning MLLMs, claiming SOTA results on MER-UniBench, MME-Emotion, and MEP-Bench.

  9. AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    AVOC is a retrieval-inspired token compression framework that improves long-form audio-video understanding in multimodal LLMs by selecting informative tokens based on classical IR principles.

  10. See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    SWIM aligns cross-attention maps from object nouns to ground-truth masks during training on the new NL-Refer dataset to enable text-only fine-grained video object understanding in MLLMs.

  11. AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    AVRT transfers reasoning to audio-visual models by distilling traces from single-modality teachers via LLM merger followed by SFT cold-start and RL, achieving SOTA on OmniBench, DailyOmni, and MMAR with 3B/7B models.

  12. C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis

    cs.CL 2026-03 unverdicted novelty 6.0 of 10

    C2F-Thinker combines structured coarse-to-fine chain-of-thought reasoning with hint-guided GRPO reinforcement learning to achieve competitive fine-grained sentiment regression and superior cross-domain generalization ...

  13. Training-Free Multimodal Large Language Model Orchestration

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    LLM Orchestration integrates modality experts via an LLM controller, cross-modal memory, and interaction layer to enable multimodal input-output without gradient-based training.

  14. Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Benchmarks seven open-source audio-video-text MLLMs on six affective datasets and shows a generative-knowledge prompting step improves fine-tuned emotion recognition.

  15. Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A new large benchmark for AI-generated human-centric video quality with pairwise preferences, plus a Mixture-of-Experts MLLM that outperforms prior methods on rating, comparison, and Q&A.

  16. CogniRoute: Learning to Route Social Evidence in Omni-Modal Models

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    CogniRoute adds a cognitive schema and route-aware RL to an omni-modal MoE, reaching 59.38% accuracy on a new 118K-example social video QA benchmark and beating prior baselines by 15-27 points.

  17. Resonant Minds: Closed-Loop Social Avatars with Theory of Mind

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    A dual-agent closed-loop system integrates Theory of Mind reasoning with multimodal video generation to create social avatars that outperform full-information baselines on dialogue quality under information asymmetry.

  18. Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective

    cs.LG 2026-04 unverdicted novelty 5.0 of 10

    CmIR uses causal inference to separate invariant causal representations from spurious ones in multimodal data, improving generalization under distribution shifts and noise via invariance, mutual information, and recon...

  19. Training-Free Multimodal Large Language Model Orchestration

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    A training-free orchestration framework integrates off-the-shelf modality experts via an LLM controller, text-centric cross-modal memory, and unified interaction layer to enable multimodal input-output without joint training.

  20. Advancing the Foundation Model for Music Understanding

    cs.SD 2025-08 unverdicted novelty 5.0 of 10

    MuFun is proposed as a unified music foundation model that jointly handles instrumental and lyrical content, and it is claimed to outperform existing audio language models on the authors' new MuCUE benchmark.

  21. FaceLLM: A Multimodal Large Language Model for Face Understanding

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Fine-tuning InternVL3 on ChatGPT-generated face QA pairs yields a face-specialized MLLM with the highest reported accuracy among MLLMs on FaceXBench.

  22. LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding

    cs.CV 2025-01 unverdicted novelty 5.0 of 10

    LLaVA-Octopus introduces instruction-driven adaptive fusion of multiple visual projectors in a multimodal LLM to improve video understanding performance.

  23. Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition

    cs.CV 2026-05 unverdicted novelty 4.0 of 10

    Rank-aware selective fusion via attention-based gating and decoupled presence/salience heads with unsupervised domain adaptation outperforms baselines and ranks 2nd on the BlEmoRE challenge for blended emotion recognition.

  24. Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition

    cs.CV 2026-05 unverdicted novelty 3.0 of 10

    A rank-aware selective fusion framework with attention gating and decoupled heads outperforms baselines and ranks 2nd on the BlEmoRE challenge for blended emotion recognition.

Pith tools