Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:40:55.394447Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 30 inbound Pith citation observations for arXiv:2501.15111.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:40:55.394447Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:28:37.058354Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:20:06.340081Z
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 97cee59c-2540-4f07-ad5f-cb50986e9527 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 352f3300-f3cd-44d2-88a6-a541c4ea4620 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70392ba5-c7c4-47f2-8574-8e3892857c3d · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Claude-3.5, 2024
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dc6adceb-c3cc-497f-8068-e8db052bbb25 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Ardila, M
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b5b6c053-5e7f-4436-99e7-8cf662e9a966 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding wav2vec 2.0: A framework for self-supervised learning of speech representations
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9f0b1b95-91c9-4d0e-8c6e-06143f7f07cb · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae54e4d2-d2fa-43ae-ab5a-ae328a9649b5 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Vggsound: A large-scale audio-visual dataset
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9bc1d275-3d71-404b-9f56-69b217851a25 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de743046-e7ae-48a9-87c3-59bfa99d8306 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 943b601f-b29d-462f-9b42-69ad4407390e · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d65ec872-4ff3-4be5-ac31-17e445cdcc9c · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 175b788c-0780-488c-bb4a-d4e750bfee86 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4872662f-d179-4211-a445-0e94321f53a5 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Qwen2-Audio Technical Report
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5637577b-5eb4-4af3-aed5-ce92a2e3cbf4 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df18e89c-f649-45a8-9c8c-b004d381ae09 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f175767c-f147-4328-8552-b51680313a66 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ad67b38-507a-4932-9a0b-a80be510e8b9 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01419a41-919b-4713-af5c-ad9aba903846 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Vita: Towards open-source interactive omni multimodal llm, 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8e7c649d-0b63-479e-9525-f3d037b92151 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Dfew: A large-scale database for recognizing dynamic facial expressions in the wild
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0a39a148-4c81-4e44-85e5-5eac0b0e6c17 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bb8be44-eceb-4d68-88ae-fde9648f8e22 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Context- aware emotion recognition networks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2e1e059b-5a48-465e-9a26-bca73e49d6ed · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53d7b1d6-d0b3-41a9-8f8d-0ae6458b32f5 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e424d400-3e77-4eda-ac64-2eb791a49cbe · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding LLaV A-med: Training a large language-and-vision assistant for biomedicine in one day
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fabd24d6-5e9c-487c-ab18-8415e1e5138b · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b1dd5db-994e-46ee-a79a-44ee7d9636a0 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 362a1ad4-03a7-458e-9c57-14d88638424b · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Mvbench: A comprehensive multi-modal video understanding benchmark, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f391b5cf-7266-43f7-baa3-99f646b96078 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Baichuan-omni technical report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 405d3d10-40c0-4371-a126-56e205d2e3cf · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Llama-vid: An image is worth 2 tokens in large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a8ceb299-0278-43b9-93eb-2962298beff4 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e1df5900-db72-480c-a907-94626c7d87fb · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding ST-LLM: Large Language Models Are Effective Temporal Learners
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74d64796-41d4-4d3b-8003-507ef3d35cf2 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding MAFW: A Large-scale, Multi-modal, Compound Affective Database for Dynamic Facial Expression Recognition in the Wild
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f5d14e7f-84f7-4876-b759-c49fcb9189a5 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69ffe85d-6394-426c-b11a-d6a0a41f1ba9 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4a474cf0-88ab-46e1-ba4a-882432e1ebd7 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0560342e-fc3c-41cd-a418-391b23b26c13 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Gpt-4 technical report, 2023
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d82264f3-547f-470c-a45e-5f5d22fcdf8f · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Gpt-4v(ision) system card
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 31501d77-30e8-4104-802f-2d58fa30ff4d · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Gpt-4o system card, 2024
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f0699602-1daa-41e3-a256-1081464faf33 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Librispeech: An asr corpus based on public domain audio books
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f4cf51d8-f745-48ff-8c4a-14b6a127afd0 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Robust speech recognition via large-scale weak supervision, 2022
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d4f2dc96-d705-46d9-80f1-a28e34abcd6f · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Mae-dfer: Efficient masked autoencoder for self-supervised dynamic facial expression recognition
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 841206ea-90a1-4f72-a555-f83656009acf · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1aa56b7e-49f1-4ffa-bf43-703d66075ded · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Gemini: A family of highly capable multimodal models, 2024
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1776cf3-950b-469c-987f-2726c427d598 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Qwen2-vl
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b6eb466f-6739-46fe-bd72-af6258ba06c2 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Qwen2.5: A party of foundation models, September 2024
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d6484f2f-be87-427d-a831-709eb4e96dcc · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding A comparison of discrete and soft speech units for improved voice conversion
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bee1f647-b8e3-404f-ae52-313aa42ee7bf · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Tarsier: Recipes for training and evaluating large video description models, 2024
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation af5ca671-aed2-4581-adf3-b18933ff2cce · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Ferv39k: A large-scale multi-scene dataset for facial expression recognition in videos, 2022
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9743f64e-6840-4f2c-b0cb-2c1089d6fd8e · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9f4ffdb-e308-4cd1-b023-ed2420731452 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Videollamb: Long video understanding with recurrent memory bridges
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ce6fb1fb-1971-4c2b-8607-808a8f8d2b09 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Mini-omni: Language models can hear, talk while thinking in streaming, 2024
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 86cb31d6-2ea3-4c92-be42-618bc0fe245c · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a66e6901-048c-406b-92c0-bbb3509f8960 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation efd17620-7e3c-4018-bdd7-77937a0a6cf7 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feab9d69-2880-4286-b13c-553afa35543b · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ea80880-3659-46dd-b108-2759b8ba247d · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6a15c7e2-f5a9-49c9-8105-46e3c39a333e · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Sigmoid loss for language image pre-training
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77a106d6-bfd1-4f3f-bb06-e63949e92388 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 57103370-9316-4edb-ad4d-a21505e2a8d8 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62bdc033-4193-46d2-9b07-93013461b598 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afe775c8-d173-4204-9530-b28bbe32f1fd · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9660cc03-9cf3-459c-8e56-385ee18c34a6 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Llava-next: A strong zero-shot video understanding model, April 2024
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7cda4ca6-c3ed-4b7f-a145-05eefa52a8a4 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Facial dynamics in video: Instruction tuning for improved facial expression perception and contextual awareness, 2025
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 34f85e75-207f-4df9-a8be-d80ba16890ec · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Llava-octopus: Unlocking instruction-driven adaptive projector fusion for video understanding, 2025
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 13d878f2-0159-490b-927c-4a1106e4c64e · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Prompting visual-language models for dynamic facial expression recognition
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e63ca3f1-2de1-4d9b-99bf-8603d6d9aaa2 · outbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e9f303b-77b4-4360-9f91-55a7dd733807 · inbound
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b08087d6-1adb-4d80-8c93-60d602ae4690 · inbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a2b2aea-b7d0-4516-80ab-a56c46060c39 · inbound
From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6abf601-e5a5-40b6-998a-9d42fbc6d159 · inbound
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fb4ef1b-8bbc-4fc1-86c2-1b82f69ac1de · inbound
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 111
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d8ab3f4-bcb6-4714-a34a-933d1bc3c8e0 · inbound
Grounding Intelligence in Movement HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 104
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e079e223-3733-428a-8926-77c04bf46313 · inbound
FaceLLM: A Multimodal Large Language Model for Face Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 100cee19-04d4-461f-b53b-523d9ca1c7ac · inbound
Advancing the Foundation Model for Music Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa56c0e6-2f13-4c61-8dc4-f259ad81b243 · inbound
Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4402c2e6-7863-41c9-b114-689858367479 · inbound
Training-Free Multimodal Large Language Model Orchestration HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 909bd532-5a03-4c10-b716-673949aa9b66 · inbound
Training-Free Multimodal Large Language Model Orchestration HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c4ad3f4d-1bc9-4903-8217-e8c010c4c8f0 · inbound
MAPF-World: Action World Model for Multi-Agent Path Finding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11cec45c-71e6-466f-97e3-23c5c3c0950d · inbound
C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4062d7e7-479f-4d9f-a260-29c62f8a83aa · inbound
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 95265503-a77e-4605-b5a6-f35c0d1754c7 · inbound
Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3f0a8955-3999-4a65-b62d-a7cc1225d133 · inbound
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a746a167-4cf3-491d-857e-f4a5e5deb2b7 · inbound
Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 283
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5bcc8013-624b-4b6c-b3b1-2f639fee0d58 · inbound
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0c2ce232-45d3-46e7-92f1-58edd413481d · inbound
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fa6cc457-39d6-4973-b9a0-195c6e5ec81d · inbound
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4ff8a9ec-fe22-4656-9a19-5af71dfaf5c6 · inbound
Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 430b824a-7ecb-463d-b4af-de283cf666fe · inbound
Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 85f127db-a05d-4fc2-868a-471b1e4a238a · inbound
Resonant Minds: Closed-Loop Social Avatars with Theory of Mind HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2b0317bc-baa6-490e-815f-2d0073361aed · inbound
AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 69e6726e-d2c1-4fdf-91df-1ca81034a101 · inbound
CogniRoute: Learning to Route Social Evidence in Omni-Modal Models HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a89b4cd4-0ccc-4af1-9a1c-c3661f978015 · inbound
AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8789fd0b-e0c6-4ee2-b437-9aeea8698f35 · inbound
Omni-Perception Policy Optimization for Multimodal Emotion Reasoning HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9152cc76-dd74-4f33-b421-a8b6edea5666 · inbound
InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d47f81cc-cea3-4ad9-a62b-73d38c8f2d38 · inbound
Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a7ecb4c-0ab1-4970-a083-caf716f08459 · inbound
E$^3$mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.