Pith. sign in

Paper Citation Record · LEDGER

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

As of 22 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 30 inbound Pith citation observations for arXiv:2501.15111.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15111 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:40:55.394447Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:28:37.058354Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:06.340081Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 97cee59c-2540-4f07-ad5f-cb50986e9527 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.430718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.144389Z digest=sha256:16f3a3a07f2214b851a3e8ed42d4b73ea4a5ebbfdeabf6e872641aa89876c510

Observation 352f3300-f3cd-44d2-88a6-a541c4ea4620 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.148677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.148677Z digest=sha256:f67b0e8401c69d9d1c2d627243cdc72bad7be22df736f1b5117ae3d5b523a628

Observation 70392ba5-c7c4-47f2-8574-8e3892857c3d · outbound

This paper cites Claude-3.5, 2024.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Claude-3.5, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.417921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.153139Z digest=sha256:3ad51ed1234614617a687313ea62328c56252b8d7099076474e0c45af922014b

Observation dc6adceb-c3cc-497f-8068-e8db052bbb25 · outbound

This paper cites Ardila, M.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Ardila, M

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.406699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.156974Z digest=sha256:11c0cf85c252f40d807f56acd02dc251924d29920169d023f76b2361c46ad4b9

Observation b5b6c053-5e7f-4436-99e7-8cf662e9a966 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.395477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.160550Z digest=sha256:8ceda09b9493bbf72cfae85c496d9ba0a9d982070783cd04e792b31c979381d4

Observation 9f0b1b95-91c9-4d0e-8c6e-06143f7f07cb · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.164421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.164421Z digest=sha256:388cc91dc67583b0df5d561f0b63fe3adb08741ba4b95beaf24ae2109bd34759

Observation ae54e4d2-d2fa-43ae-ab5a-ae328a9649b5 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Vggsound: A large-scale audio-visual dataset

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.384033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.168586Z digest=sha256:f05b8f61abda49b13a3d3f6a7c4a7e6886cc8171b2a46976fea0689c9d59549d

Observation 9bc1d275-3d71-404b-9f56-69b217851a25 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.172325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.172325Z digest=sha256:6ed07fd7a2ed4b6bc55c39b483d15dc564f687e326b3651ef702e3b218f99eb9

Observation de743046-e7ae-48a9-87c3-59bfa99d8306 · outbound

This paper cites MotionLLM: Understanding Human Behaviors from Human Motions and Videos.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.176113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.176113Z digest=sha256:df2776e29c6939b9431b4526b0fb9f1ddc72729b62bab3b669009f8f2ec25669

Observation 943b601f-b29d-462f-9b42-69ad4407390e · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.372350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.179856Z digest=sha256:36e1a19a95837e953e3d5ffbc4051ba41557726b58e029134b78c8c6722ecaa4

Observation d65ec872-4ff3-4be5-ac31-17e445cdcc9c · outbound

This paper cites Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.184221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.184221Z digest=sha256:18c0043ac78054a8c7b01c2e3c443f4104b6ac78837eaaca374a68b330725905

Observation 175b788c-0780-488c-bb4a-d4e750bfee86 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.188011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.188011Z digest=sha256:f0f1a4415330b0816978ac25c77ed6fcfc1dc2317da02a47452980d01f38f0a5

Observation 4872662f-d179-4211-a445-0e94321f53a5 · outbound

This paper cites Qwen2-Audio Technical Report.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Qwen2-Audio Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.192173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.192173Z digest=sha256:b851f158e96af68683ada1c4be30cd0b9b1023f5470882d648f6a1aea43a45ae

Observation 5637577b-5eb4-4af3-aed5-ce92a2e3cbf4 · outbound

This paper cites FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.195830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.195830Z digest=sha256:be172f8c616227649fd2ecb0c8455ecfa6bc9731f7f0d5d6a6b0b4bd43df6087

Observation df18e89c-f649-45a8-9c8c-b004d381ae09 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.199576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.199576Z digest=sha256:8bf617939087ab581f720c9201117dad035fa90b0a5c264485da57d6145d525c

Observation f175767c-f147-4328-8552-b51680313a66 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.203595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.203595Z digest=sha256:306dcde6e6d42d1eefc47b69019fd7193f1cad412d4342f51b570b854b8dfa63

Observation 5ad67b38-507a-4932-9a0b-a80be510e8b9 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.207606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.207606Z digest=sha256:231a4ea0bcb6687a14ddb7102213d3e0548e4f4b2d96552a17dbd98506fcda5f

Observation 01419a41-919b-4713-af5c-ad9aba903846 · outbound

This paper cites Vita: Towards open-source interactive omni multimodal llm, 2024.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Vita: Towards open-source interactive omni multimodal llm, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.359730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.211317Z digest=sha256:4fee6bdb9baac6b9f39972f9aca81bb70abfcf1737ce296902fd933e9d36bd61

Observation 8e7c649d-0b63-479e-9525-f3d037b92151 · outbound

This paper cites Dfew: A large-scale database for recognizing dynamic facial expressions in the wild.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Dfew: A large-scale database for recognizing dynamic facial expressions in the wild

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.345488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.214771Z digest=sha256:898d603da522d1da675df88a0b820779d822ad49764e268d8899dc1c8c23502a

Observation 0a39a148-4c81-4e44-85e5-5eac0b0e6c17 · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.218443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.218443Z digest=sha256:f1dbadd483aaf132853d7b179a9603e30c608d208b040dd075b2d3918114717a

Observation 6bb8be44-eceb-4d68-88ae-fde9648f8e22 · outbound

This paper cites Context- aware emotion recognition networks.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Context- aware emotion recognition networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.331739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.222118Z digest=sha256:7650808761ea8887382e5b7f06e83d44f8ed838d7059637cf8ccb9802f2980b7

Observation 2e1e059b-5a48-465e-9a26-bca73e49d6ed · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.225650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.225650Z digest=sha256:e9f41b4a0b3a8937adb0548f8d088c474f7a65ab320582b11b3b34f3966a1b36

Observation 53d7b1d6-d0b3-41a9-8f8d-0ae6458b32f5 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.229080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.229080Z digest=sha256:dca10486995c43964de059117ed50527b658d0a4d87a978edaa4b4502af69e7c

Observation e424d400-3e77-4eda-ac64-2eb791a49cbe · outbound

This paper cites LLaV A-med: Training a large language-and-vision assistant for biomedicine in one day.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding LLaV A-med: Training a large language-and-vision assistant for biomedicine in one day

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.315481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.232986Z digest=sha256:ad6e7dbcb8a2aaf9ee3a204f32a68628484f1e05c32e12981f71094301a1c499

Observation fabd24d6-5e9c-487c-ab18-8415e1e5138b · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.236516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.236516Z digest=sha256:0baffcd307e97ef72a99824d5e45e594b236df0b1d0af58cc422138a0ce607fa

Observation 6b1dd5db-994e-46ee-a79a-44ee7d9636a0 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.240225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.240225Z digest=sha256:e2cc2e00679127566c3923f2e4292ad574d1ea28a6b0478b17c071012f584946

Observation 362a1ad4-03a7-458e-9c57-14d88638424b · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark, 2024.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Mvbench: A comprehensive multi-modal video understanding benchmark, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.300713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.244387Z digest=sha256:32e3bf5be130d85d2ed64aa8e37a416056199dc4260d7afa1efbd55f66b85c03

Observation f391b5cf-7266-43f7-baa3-99f646b96078 · outbound

This paper cites Baichuan-omni technical report.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Baichuan-omni technical report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.248180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.248180Z digest=sha256:5361fc0defd67d063f87a4ec5d3c19bbffe54e99aa5a681f04759db81826878b

Observation 405d3d10-40c0-4371-a126-56e205d2e3cf · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Llama-vid: An image is worth 2 tokens in large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.285755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.252200Z digest=sha256:668baffbf4e8ca2311eea55632ac2f637de2921575069dd24cc9527e35152815

Observation a8ceb299-0278-43b9-93eb-2962298beff4 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.271932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.255847Z digest=sha256:5d1430aad0b1255391c378ca0d04852c61718c371e14a19a64721f0688ffe789

Observation e1df5900-db72-480c-a907-94626c7d87fb · outbound

This paper cites ST-LLM: Large Language Models Are Effective Temporal Learners.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding ST-LLM: Large Language Models Are Effective Temporal Learners

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.259514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.259514Z digest=sha256:ea2c5950a7d39f81610895c23ae1f317729982eb7f48a04261c2531cd5b7b2d6

Observation 74d64796-41d4-4d3b-8003-507ef3d35cf2 · outbound

This paper cites MAFW: A Large-scale, Multi-modal, Compound Affective Database for Dynamic Facial Expression Recognition in the Wild.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding MAFW: A Large-scale, Multi-modal, Compound Affective Database for Dynamic Facial Expression Recognition in the Wild

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.257099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.263436Z digest=sha256:9363497ec64b2261ca4b705dca8c1f5e7a31906ccf194250607f749977d0fd2f

Observation f5d14e7f-84f7-4876-b759-c49fcb9189a5 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.266841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.266841Z digest=sha256:644477928c57cc6db8484f38fc175da2c5f5b6c0daefb25cf3dcceada41e6668

Observation 69ffe85d-6394-426c-b11a-d6a0a41f1ba9 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.244134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.271144Z digest=sha256:b1ea8f9aac561d1fb342e78ab55bed08879f309c7200c62523edf1805630498a

Observation 4a474cf0-88ab-46e1-ba4a-882432e1ebd7 · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.230620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.275425Z digest=sha256:1b2d042329f18825c2e2743858b91a27e14a2681c1e89aea4656dd02f1315a50

Observation 0560342e-fc3c-41cd-a418-391b23b26c13 · outbound

This paper cites Gpt-4 technical report, 2023.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Gpt-4 technical report, 2023

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.279023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.279023Z digest=sha256:e574f2c4de566668c6fabbb55bd25baec5739c2725b551979dc67b76da17079b

Observation d82264f3-547f-470c-a45e-5f5d22fcdf8f · outbound

This paper cites Gpt-4v(ision) system card.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Gpt-4v(ision) system card

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.207989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.282637Z digest=sha256:d518f1e78ebb663de33bb2d4f9c7e878e741c11157e657e483ac93d0f9148230

Observation 31501d77-30e8-4104-802f-2d58fa30ff4d · outbound

This paper cites Gpt-4o system card, 2024.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Gpt-4o system card, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.069842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.286067Z digest=sha256:efe69652c04402d2d22d55737bcc7beb14245ef476ee03bb8bc96b126239fc48

Observation f0699602-1daa-41e3-a256-1081464faf33 · outbound

This paper cites Librispeech: An asr corpus based on public domain audio books.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Librispeech: An asr corpus based on public domain audio books

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.058737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.289506Z digest=sha256:cea7c4ae76b990ce05c7f5b86b096251f590b67a94ef1cf2236703f58c85a7f0

Observation f4cf51d8-f745-48ff-8c4a-14b6a127afd0 · outbound

This paper cites Robust speech recognition via large-scale weak supervision, 2022.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Robust speech recognition via large-scale weak supervision, 2022

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.046151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.293365Z digest=sha256:d0e1feff995a38059a65466f86eac32fc22e6c0f7a618cb677126d3d1a61122b

Observation d4f2dc96-d705-46d9-80f1-a28e34abcd6f · outbound

This paper cites Mae-dfer: Efficient masked autoencoder for self-supervised dynamic facial expression recognition.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Mae-dfer: Efficient masked autoencoder for self-supervised dynamic facial expression recognition

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.031883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.296928Z digest=sha256:6a55c123ba3a20f486fc4a5bee368a112bf80e6265c03a5a9c636b52637bf4f5

Observation 841206ea-90a1-4f72-a555-f83656009acf · outbound

This paper cites HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-10T14:40:55.515579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.300372Z digest=sha256:de778d27a92d4b49abea9b327d7e5b38e45d96d748f4799498065316d831a622

Observation 1aa56b7e-49f1-4ffa-bf43-703d66075ded · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2024.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Gemini: A family of highly capable multimodal models, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.304374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.304374Z digest=sha256:bfd7642af092516286b279f419c34ac425c6df5a468b818654cf84749ef7903e

Observation c1776cf3-950b-469c-987f-2726c427d598 · outbound

This paper cites Qwen2-vl.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Qwen2-vl

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:56.010703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.308080Z digest=sha256:ad64f150488466e6405d13acde5a66dfa3e9000b75a489138d961068dcd2a5d7

Observation b6eb466f-6739-46fe-bd72-af6258ba06c2 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Qwen2.5: A party of foundation models, September 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:55.997146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.311694Z digest=sha256:29925b48c782eb4f9ab8bc47cbb6e9cd0cadb2089a0c05a466bfba4a8a922803

Observation d6484f2f-be87-427d-a831-709eb4e96dcc · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding A comparison of discrete and soft speech units for improved voice conversion

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:55.984674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.315479Z digest=sha256:b1a9e680bed2aa2c63fdb321cfaaa6f0e746a91930c099fc8e7e7febcde871a8

Observation bee1f647-b8e3-404f-ae52-313aa42ee7bf · outbound

This paper cites Tarsier: Recipes for training and evaluating large video description models, 2024.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Tarsier: Recipes for training and evaluating large video description models, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:55.970787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.319212Z digest=sha256:fedab53a4ce5424bd29c8478acb5404cb30c7b86f7ba8cdecb50f9f672e1bcb9

Observation af5ca671-aed2-4581-adf3-b18933ff2cce · outbound

This paper cites Ferv39k: A large-scale multi-scene dataset for facial expression recognition in videos, 2022.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Ferv39k: A large-scale multi-scene dataset for facial expression recognition in videos, 2022

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:55.956828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.322965Z digest=sha256:1f2b59576b4b00d857b189401fa5d1e4a65ad0f6dec7a5cbeec8cc3446725973

Observation 9743f64e-6840-4f2c-b0cb-2c1089d6fd8e · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.326472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.326472Z digest=sha256:5382af972711876085aa299617b5fa140fe96d262ae90835b8612097290859ff

Observation d9f4ffdb-e308-4cd1-b023-ed2420731452 · outbound

This paper cites Videollamb: Long video understanding with recurrent memory bridges.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Videollamb: Long video understanding with recurrent memory bridges

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:55.945653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.330479Z digest=sha256:8be41d0ab39eb6c83baec18f617f31a141fc22a94cf246694bf1ecadb725dd51

Observation ce6fb1fb-1971-4c2b-8607-808a8f8d2b09 · outbound

This paper cites Mini-omni: Language models can hear, talk while thinking in streaming, 2024.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Mini-omni: Language models can hear, talk while thinking in streaming, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:55.932936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.334021Z digest=sha256:045a58ae8faaf5a5c288b822a3ac11f002c5868ac46b93d3ff964823c3df27c6

Observation 86cb31d6-2ea3-4c92-be42-618bc0fe245c · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.337372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.337372Z digest=sha256:0f2326087d0662b9aca2cd6887ed179a48de90f01e73f2b9e927f1d43ba8c574

Observation a66e6901-048c-406b-92c0-bbb3509f8960 · outbound

This paper cites Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:55.919573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.341090Z digest=sha256:3b0204c383dfcde04658332f0e0506ed5bf082dd28192e998e5d52d8805b8fa8

Observation efd17620-7e3c-4018-bdd7-77937a0a6cf7 · outbound

This paper cites Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.344331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.344331Z digest=sha256:001b8ee0a04a0972cb358e469b7841de610cead047806eabfc5d2fc794e100c6

Observation feab9d69-2880-4286-b13c-553afa35543b · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.351904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.351904Z digest=sha256:54a957c74378cc6b12acb5506c100003d715586bd8b1f72f9cb6f5513addf449

Observation 2ea80880-3659-46dd-b108-2759b8ba247d · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:55.906057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.355767Z digest=sha256:250660743c249d7c57bdd1b889316b85ccafa1d0eae56bf974edae68283018fb

Observation 6a15c7e2-f5a9-49c9-8105-46e3c39a333e · outbound

This paper cites Sigmoid loss for language image pre-training.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Sigmoid loss for language image pre-training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.359554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.359554Z digest=sha256:07648a8099695924c48097ca1dc808574a6d63d488758e0dd5f7273eb3c5f39e

Observation 77a106d6-bfd1-4f3f-bb06-e63949e92388 · outbound

This paper cites Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:55.885164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.363524Z digest=sha256:3ea14e11a6070e5f6b449c4efc96681cff99cfe038205bf49292ea5af17dd089

Observation 57103370-9316-4edb-ad4d-a21505e2a8d8 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.368513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.368513Z digest=sha256:50a5bb4ad45a4e49e49280453d305bad5514e6b4cb73a229b767f8235bf2ba08

Observation 62bdc033-4193-46d2-9b07-93013461b598 · outbound

This paper cites InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.372307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.372307Z digest=sha256:be051c8834e29bf7b271f7a5d0f7d12f7a769e9ee73b833d7f7d516b5d6291b0

Observation afe775c8-d173-4204-9530-b28bbe32f1fd · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.376225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.376225Z digest=sha256:76a05bf2c69146d3c2a44738123ddb008257d0f8ce0e03e85dd0142bdc668f05

Observation 9660cc03-9cf3-459c-8e56-385ee18c34a6 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, April 2024.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Llava-next: A strong zero-shot video understanding model, April 2024

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:55.872405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.379728Z digest=sha256:c4665aaa4f42878d92f3f1edcbf972b9589314086ea0025ee1983ece88cdfd25

Observation 7cda4ca6-c3ed-4b7f-a145-05eefa52a8a4 · outbound

This paper cites Facial dynamics in video: Instruction tuning for improved facial expression perception and contextual awareness, 2025.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Facial dynamics in video: Instruction tuning for improved facial expression perception and contextual awareness, 2025

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:55.860420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.383221Z digest=sha256:f31beda762f411ea35e1b244947e8896e61f7b3de37464e821cc9baa8d1c9331

Observation 34f85e75-207f-4df9-a8be-d80ba16890ec · outbound

This paper cites Llava-octopus: Unlocking instruction-driven adaptive projector fusion for video understanding, 2025.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Llava-octopus: Unlocking instruction-driven adaptive projector fusion for video understanding, 2025

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:55.846580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.386725Z digest=sha256:2c7112e9d083ca2f289d32faf50acb533737832276ff3a43a227b3c89981b428

Observation 13d878f2-0159-490b-927c-4a1106e4c64e · outbound

This paper cites Prompting visual-language models for dynamic facial expression recognition.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Prompting visual-language models for dynamic facial expression recognition

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:40:55.831514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:40:55.390431Z digest=sha256:28c4a2913949a63564e0ba1b2df958ad7f1a2fdd89a4129f38c611cb63cd84b2

Observation e63ca3f1-2de1-4d9b-99bf-8603d6d9aaa2 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.394447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.394447Z digest=sha256:5bafe32372864e83666c7f1f0b83cadf923a850d93034ca043d5050a3eeb45a7

Pith citing papers

Observation 0e9f303b-77b4-4360-9f91-55a7dd733807 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.563253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:c1aa7b2f4ded5e8e4dc635e3d7b79c5f9e92b9b29a235a98a03d205e5eb103e8

Observation b08087d6-1adb-4d80-8c93-60d602ae4690 · inbound

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding cites this paper.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:37.058354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:37.058354Z digest=sha256:4c87a90396fca206c8415bf3f3cccbae5551da1833661f80be1b70b57213d4bb

Observation 6a2b2aea-b7d0-4516-80ab-a56c46060c39 · inbound

From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models cites this paper.

From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:19.930713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:19.930713Z digest=sha256:ba3b4b683358e5d166ebfd5f62cd3796c9688633fd65ebac989e06cdeeef98be

Observation b6abf601-e5a5-40b6-998a-9d42fbc6d159 · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.070930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.070930Z digest=sha256:8e9522e87081014bbcfeb46b58fe1849e75e3b5488ad3c8fca05fb36c15de06f

Observation 1fb4ef1b-8bbc-4fc1-86c2-1b82f69ac1de · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:36.656003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:36.656003Z digest=sha256:574b4ce9aabc008fa810f268b1469e41b42022d540e20ec418e519440dbe5028

Observation 2d8ab3f4-bcb6-4714-a34a-933d1bc3c8e0 · inbound

Grounding Intelligence in Movement cites this paper.

Grounding Intelligence in Movement HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:32.881282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:32.881282Z digest=sha256:7b23ab2579f3aa1f60d80597060b021992ffca21aa8fb6416d59ff46f514b7b4

Observation e079e223-3733-428a-8926-77c04bf46313 · inbound

FaceLLM: A Multimodal Large Language Model for Face Understanding cites this paper.

FaceLLM: A Multimodal Large Language Model for Face Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T17:39:41.688863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:39:41.688863Z digest=sha256:7367c5f2d6f64b76a7ff6a71c0c759db66e8638a59393dd7a609319c64776eb6

Observation 100cee19-04d4-461f-b53b-523d9ca1c7ac · inbound

Advancing the Foundation Model for Music Understanding cites this paper.

Advancing the Foundation Model for Music Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T05:49:55.526816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:49:55.526816Z digest=sha256:3fd6cebf3ff425c3516b56e52c1c96c5841f4673061443cd53c33f0c011bfb11

Observation fa56c0e6-2f13-4c61-8dc4-f259ad81b243 · inbound

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting cites this paper.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.540778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.540778Z digest=sha256:0d699dde136d171078d1a85f8267bb80a4ae9510ac023642f24f5171c76c5c49

Observation 4402c2e6-7863-41c9-b114-689858367479 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:12:54.107948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:f5a3b80c5f70b1454371b0ed48462ead393bd7915d05570a55cc5afff8c09066

Observation 909bd532-5a03-4c10-b716-673949aa9b66 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:05:30.643788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:700afc869989609846550f2a8165fe7a86ce7b7d80ee4e3af11d8f37ee6546a0

Observation c4ad3f4d-1bc9-4903-8217-e8c010c4c8f0 · inbound

MAPF-World: Action World Model for Multi-Agent Path Finding cites this paper.

MAPF-World: Action World Model for Multi-Agent Path Finding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:28:24.016809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:28:24.016809Z digest=sha256:a9c20ead78db47030cebaec8a2060c94a1e0f2cca9ee3b7f995ee82bdc214845

Observation 11cec45c-71e6-466f-97e3-23c5c3c0950d · inbound

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis cites this paper.

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:55:53.282103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T13:51:40.334057Z digest=sha256:e673e183dd46213797a1baefc54bb0155279abfc70a129fc43512ac3d65ee985

Observation 4062d7e7-479f-4d9f-a260-29c62f8a83aa · inbound

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration cites this paper.

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:48.970344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T20:20:08.590407Z digest=sha256:1699bd0d4d868705d755fee789ccb0971b865fbfa7ec3993f81c5d01d670026b

Observation 95265503-a77e-4605-b5a6-f35c0d1754c7 · inbound

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions cites this paper.

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:53:04.276506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T08:49:33.107658Z digest=sha256:57f570f5896dd239154f6797526517c0e92c1c8ada422be1e04ce3f205e314ae

Observation 3f0a8955-3999-4a65-b62d-a7cc1225d133 · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:18:32.134430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:6fbab28fbcc5436b80d2e09ac66ffe55e75bef29356e196cae51afc47312bd03

Observation a746a167-4cf3-491d-857e-f4a5e5deb2b7 · inbound

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective cites this paper.

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 283

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:51:03.220651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T04:32:29.428080Z digest=sha256:ff27d6d2f477a91afeb837e8a82961ecbc81d7bd1dcb72309e6229a2b2a6fdee

Observation 5bcc8013-624b-4b6c-b3b1-2f639fee0d58 · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:56.047112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:41632fcdd2fc19e10e4efb4638bc356a2d3c0b14371e24de748877e9afb0b211

Observation 0c2ce232-45d3-46e7-92f1-58edd413481d · inbound

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions cites this paper.

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 91

Resolution
malformed identifier
arxiv_id, observed 2026-05-20T18:53:39.002635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T18:49:18.815456Z digest=sha256:f21d26b36aea57d83e1e6fce006f72e6885d531723d822005c8fae6635010bac

Observation fa6cc457-39d6-4973-b9a0-195c6e5ec81d · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.342785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:683118350b4af57a078d7a0ae84acc6554281d0009892deaea5248e5ca81651f

Observation 4ff8a9ec-fe22-4656-9a19-5af71dfaf5c6 · inbound

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition cites this paper.

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:19:39.200505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:18:29.720630Z digest=sha256:f54cf5f05c62abb72659c578ac6def6f6d531accf07fcc5ef551b60335340084

Observation 430b824a-7ecb-463d-b4af-de283cf666fe · inbound

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition cites this paper.

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:14:57.327465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T17:06:26.561698Z digest=sha256:b124a02d3c3d894fa0506c84b4a8d864375086c19b7220dd1c4b83316c57440c

Observation 85f127db-a05d-4fc2-868a-471b1e4a238a · inbound

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind cites this paper.

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:56.204673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T02:04:39.753443Z digest=sha256:ff9a84d6e5b713a1832dd64800a86e3955744245eb07da9e30bfe2579c12718c

Observation 2b0317bc-baa6-490e-815f-2d0073361aed · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.372188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:d7dc8de3b6343617e2a6ddb9e30ff6e9d5125097954edffd3f3dd8144b5dba59

Observation 69e6726e-d2c1-4fdf-91df-1ca81034a101 · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.398553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:7b2d48f06595e1e9d0c7c442863fee8210d3df1cebf51898ab1373ad17fda2bd

Observation a89b4cd4-0ccc-4af1-9a1c-c3661f978015 · inbound

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression cites this paper.

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.283536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:16:56.174638Z digest=sha256:e414391f13acbdcbf475bb1500cf19b3c8c8cec46d8e193d66e04d5858c425e9

Observation 8789fd0b-e0c6-4ee2-b437-9aeea8698f35 · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:20:06.342275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-25T21:31:38.450382Z digest=sha256:9c691fff97f8ad23ab5db8b7c3177cd8fc0808e6e9b73b568918365491ba3af0

Observation 9152cc76-dd74-4f33-b421-a8b6edea5666 · inbound

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring cites this paper.

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:09.040422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:25:09.040422Z digest=sha256:8d506af24e17e7adc223901a3b68eac42159d10bd1965f6f21e57743f6dcaa31

Observation d47f81cc-cea3-4ad9-a62b-73d38c8f2d38 · inbound

Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model cites this paper.

Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T20:07:41.111549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:07:41.111549Z digest=sha256:35b1474a33984adb4ef894982a5d1d643fb3b1870a3f53b81a7e107f2c87307d

Observation 2a7ecb4c-0ab1-4970-a083-caf716f08459 · inbound

E$^3$mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment cites this paper.

E$^3$mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T17:17:41.350190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:17:41.350190Z digest=sha256:5be0e1d66044de3c0bb862808386d35a2f6b2b4df8679a038f16122c9281efb4