Pith. sign in

hub

Humanomni: A large vision-speech language model for human-centric video understanding

16 Pith papers cite this work. Polarity classification is still indexing.

16 Pith papers citing it

hub tools

citation-role summary

background 1 baseline 1

citation-polarity summary

years

2026 14 2025 2

representative citing papers

Training-Free Multimodal Large Language Model Orchestration

cs.CL · 2025-08-06 · unverdicted · novelty 6.0 · 2 refs

LLM Orchestration integrates modality experts via an LLM controller, cross-modal memory, and interaction layer to enable multimodal input-output without gradient-based training.

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind

cs.CV · 2026-06-04 · unverdicted · novelty 5.0

A dual-agent closed-loop system integrates Theory of Mind reasoning with multimodal video generation to create social avatars that outperform full-information baselines on dialogue quality under information asymmetry.

citing papers explorer

Showing 16 of 16 citing papers.