Pith. sign in

Paper Citation Record · LEDGER

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2412.09596.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09596 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:47:39.376946Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:29.961741Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c4b67586-0fda-4dfd-a13e-849b0c11f5c8 · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.376946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.376946Z digest=sha256:98c760377318790fb2b7dec05ac3f66cf182976373497d0d3c468969e214f15d

Observation 75b57934-9208-4775-af26-56708e043fd2 · inbound

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs cites this paper.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:27.783266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:911a6177acc922759f158f2140b123a3f7429f70554100a34c6e60625f22ad95

Observation e7381476-3277-4121-a530-b252841a80b1 · inbound

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence cites this paper.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:43.022715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:43.022715Z digest=sha256:fcef20fc89613cbfe6523955e523acbc7614cabdfb236d3d2e76c058d636e02b

Observation 0adb5e62-b5c0-4719-9158-042dcd1b604f · inbound

Know-MRI: A Knowledge Mechanisms Revealer&Interpreter for Large Language Models cites this paper.

Know-MRI: A Knowledge Mechanisms Revealer&Interpreter for Large Language Models InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:35.710503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:35.710503Z digest=sha256:222acf74da77e88b99e9633f1d38b035dc3112bdc506814cc9a7f2eecc312259

Observation 4cef6c13-6340-4263-8359-99ce2afbf1c6 · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.167998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.167998Z digest=sha256:0bb5af6d470bd2285dd62549b47b97d97a05231b017fe75085951bd26060d488

Observation 3ec31ea1-a90e-4493-a56b-c09be0b7b9bb · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.198672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.198672Z digest=sha256:ae962524eef8f48af1ddc856a3f76aa00e0f3208edc97e09948df76440f5ee03

Observation 6b46ae8c-9509-456b-aa26-197f2265ce28 · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:11.657893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:11.657893Z digest=sha256:995653a670c18634b470329045b7cdfbac911ead2e2d1ee3dab90afcc1209a2b

Observation c6146844-f8ee-49d0-9297-b77bace798b1 · inbound

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning cites this paper.

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-21T16:44:16.021335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T16:43:11.995960Z digest=sha256:5cfaef22dad58e139b9c980045e07cdfc2486d32d6a923a2611734d3ab340280

Observation 12da96cb-a329-43fa-9797-2526cd9f7f49 · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:25:30.155575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:80162292ddb1072da790271dedea7943b8c21fd622edeeed12adada98def5904

Observation e404bb10-364e-4aa7-977c-1807aaf36f69 · inbound

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games cites this paper.

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:19:13.658047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:17:02.332687Z digest=sha256:b9affcca41fbc7f5a33d53b1929c1f19f3d709496c603dafc6708850707721c3

Observation d8002c6b-4f07-46a7-b86c-10a96f7a00a3 · inbound

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference cites this paper.

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:29.964055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T18:17:53.013043Z digest=sha256:ad61d835a37c29871c48492d9df151db680cec1d423413ef3f2a9e1a698fab9d

Observation a25e7ef2-4c1f-492c-8e1e-295760706f59 · inbound

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory cites this paper.

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-11T06:35:35.951554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T06:35:35.951554Z digest=sha256:df4733ad5e0625cf73928f2d21b18e8906911f1b1f87f5dd424c3c78c94d47a0

Observation c7153a65-386a-4160-b939-63be0ae1703d · inbound

FOLIO: Focused Semantic Memory for Streaming Video Understanding cites this paper.

FOLIO: Focused Semantic Memory for Streaming Video Understanding InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:49.748388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:49.748388Z digest=sha256:4585069f3c9b7642ae06d5f0f4758ebead360c4861abd8886fc2289168de29f8

Observation 8a354390-4034-47f5-b432-5c9212b803d4 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 134

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.654651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.654651Z digest=sha256:87e7df20d6f0feb9e8fd4bdb0cdb66c613b12a45b5d2ada952986e9284d50035

Observation 23380163-05c6-4472-aa1b-ba3150efe153 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.894956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.894956Z digest=sha256:a96394fcad986d4e79002410f9755a707c21ee6d68746909ac1c05d1fdbdaf4e