Pith. sign in

Paper Citation Record · LEDGER

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2501.15111.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15111 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:55:19.930713Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:06.340081Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0e9f303b-77b4-4360-9f91-55a7dd733807 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.563253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:22c3b4341b65d5e64cc82ef0b459f815401025bc8af30246fb1461b8d99f4a66

Observation 6a2b2aea-b7d0-4516-80ab-a56c46060c39 · inbound

From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models cites this paper.

From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:19.930713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:19.930713Z digest=sha256:03a4a57f69f6fd78c771381e79d24e17ebe1ff837b7e5cd4ce78b6af555d5082

Observation b6abf601-e5a5-40b6-998a-9d42fbc6d159 · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.070930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.070930Z digest=sha256:896d495410c5ad08c1c465a61ee1a0d9639e810055634e456ad09ddcfe3194fc

Observation 1fb4ef1b-8bbc-4fc1-86c2-1b82f69ac1de · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:36.656003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:36.656003Z digest=sha256:830df917dccf67ccd36e8309a556af256cb9581619116788781a09e214d5c069

Observation 2d8ab3f4-bcb6-4714-a34a-933d1bc3c8e0 · inbound

Grounding Intelligence in Movement cites this paper.

Grounding Intelligence in Movement HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:32.881282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:32.881282Z digest=sha256:6cdc552d2d11be9cfa758a8ed7fad307057c07d2334e3ca2cc09e74b8986facd

Observation e079e223-3733-428a-8926-77c04bf46313 · inbound

FaceLLM: A Multimodal Large Language Model for Face Understanding cites this paper.

FaceLLM: A Multimodal Large Language Model for Face Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T17:39:41.688863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:39:41.688863Z digest=sha256:4bf6dc626466a68873d6348bae723f1d409e8215c9ce909e32f8881557f105fc

Observation 100cee19-04d4-461f-b53b-523d9ca1c7ac · inbound

Advancing the Foundation Model for Music Understanding cites this paper.

Advancing the Foundation Model for Music Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T05:49:55.526816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:49:55.526816Z digest=sha256:86cc8de13887d9e39dbeb9485e667d3f956fd5440aeedec10a371869b926b954

Observation fa56c0e6-2f13-4c61-8dc4-f259ad81b243 · inbound

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting cites this paper.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.540778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.540778Z digest=sha256:57f572166d56775e0370f5ecec78f733a00b7fa0f43fcf21ea412579417ce125

Observation 4402c2e6-7863-41c9-b114-689858367479 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:12:54.107948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:a591e222d3e4b20d65ed27ce0d20116901c9d38dc351836df7330716bdd763a2

Observation 909bd532-5a03-4c10-b716-673949aa9b66 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:05:30.643788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:b09ff98a1b962dbb9660f7f01f3d5b066d3f331c7abdb000be82a9e6789deb9e

Observation 11cec45c-71e6-466f-97e3-23c5c3c0950d · inbound

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis cites this paper.

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:55:53.282103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T13:51:40.334057Z digest=sha256:d3c3a2986f4595de156693a6e3f94e510130852ce5128b55278b3297322a0121

Observation 4062d7e7-479f-4d9f-a260-29c62f8a83aa · inbound

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration cites this paper.

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:48.970344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T20:20:08.590407Z digest=sha256:2d8192dcdcd1c9900795c15fa86a26473a259446eeadb0f0dbbf1cc5e9624df3

Observation 95265503-a77e-4605-b5a6-f35c0d1754c7 · inbound

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions cites this paper.

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:53:04.276506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:49:33.107658Z digest=sha256:3cd01b5906a2e6d307bdc80c05619ef95ea1eb46982c1c053d7c815498a20539

Observation 3f0a8955-3999-4a65-b62d-a7cc1225d133 · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:18:32.134430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:e1753531b9673b77024ef8ad00a32b4b5217ab252c43ecc7307889c617aac08d

Observation a746a167-4cf3-491d-857e-f4a5e5deb2b7 · inbound

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective cites this paper.

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 283

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:51:03.220651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:32:29.428080Z digest=sha256:457f59721d6845409551d65d0316bf943b63ebf1d86a054c3b54681022d13f58

Observation 5bcc8013-624b-4b6c-b3b1-2f639fee0d58 · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:56.047112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:ce2c14588c6cad0f564de1a956050ff5637db33774f94568596b7a5c98ba7d06

Observation 0c2ce232-45d3-46e7-92f1-58edd413481d · inbound

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions cites this paper.

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 91

Resolution
malformed identifier
arxiv_id, observed 2026-05-20T18:53:39.002635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T18:49:18.815456Z digest=sha256:9006eabfd70cd2deef86c824f747c30e20ff7a245a5316edac9a15482dba9b10

Observation fa6cc457-39d6-4973-b9a0-195c6e5ec81d · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.342785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:7856d5efbd35797e39c16ea9fc58f2703b649957381c459a6452a69837246ded

Observation 4ff8a9ec-fe22-4656-9a19-5af71dfaf5c6 · inbound

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition cites this paper.

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:19:39.200505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T05:18:29.720630Z digest=sha256:2ef7ae705d6133935540785ca30b68cc82a55ec8998275df502ae6b85bbe44da

Observation 430b824a-7ecb-463d-b4af-de283cf666fe · inbound

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition cites this paper.

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:14:57.327465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T17:06:26.561698Z digest=sha256:c8c726eafaea590f7f82ec3c47225b80a732d96488da664588f31abac8bc9945

Observation 85f127db-a05d-4fc2-868a-471b1e4a238a · inbound

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind cites this paper.

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:56.204673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T02:04:39.753443Z digest=sha256:391a8e991ea6ee05f5c4dc83cf87c1712528e9ea058c16e87ab71e6093541dad

Observation 2b0317bc-baa6-490e-815f-2d0073361aed · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.372188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:964ae4342b01ff1485a0c796908704e763587cdcf42735998aaeb5659a377919

Observation 69e6726e-d2c1-4fdf-91df-1ca81034a101 · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.398553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:89bb11b1a01b7b9ac765f390e3c1fbb30f32a6f9e363caa51a39f2b42f9f1411

Observation a89b4cd4-0ccc-4af1-9a1c-c3661f978015 · inbound

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression cites this paper.

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.283536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T00:16:56.174638Z digest=sha256:edeef8d30811ad97ea27f571d92c2e91d7731438419570e360c96d410ef2e039

Observation 8789fd0b-e0c6-4ee2-b437-9aeea8698f35 · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:20:06.342275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T21:31:38.450382Z digest=sha256:d7be1e3d77237a4d7d772569c8cb5e0efa68694d1deffa43b14faac621f224cb

Observation 9152cc76-dd74-4f33-b421-a8b6edea5666 · inbound

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring cites this paper.

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:09.040422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:25:09.040422Z digest=sha256:639d021d24429108ce32dd35d2fcbff8a7efa31939d896d15583580843c6b0ca

Observation d47f81cc-cea3-4ad9-a62b-73d38c8f2d38 · inbound

Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model cites this paper.

Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T20:07:41.111549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:07:41.111549Z digest=sha256:ece15bf6d07b95a27ec2577cf135a48a88d81f211e99504337e5a0e31a405ca7