Pith. sign in

Paper Citation Record · LEDGER

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams

As of 6 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 0 inbound Pith citation observations for arXiv:2606.17798.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.17798 v1

Coverage vector

measured 100 of 100 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T01:12:46.295455Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 100 outbound references displayed

  • verified exact40
  • verified fuzzy0
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 627e6066-a645-4ca2-b461-9f5cb92c2c27 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.140145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:8e938db3d0c0a881c271c16be17431f63da1ac59856f120de09a5c8ae53e2fe5

Observation 3f98e227-1dd8-493d-b85a-463a544ec722 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.173494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:71829d4484d813ae7ef013754a4172ed64bcf4303bc3dc8077c8a316c99e0e3c

Observation 0aabb6b9-16c7-4b69-b664-000368da88b4 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.208066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:87473c671309ac1488e4c568ec1f743ee6750bf9f1bd99ed11ad08a8f6bc775f

Observation 0d5f52e1-f1a5-468e-80b3-7343718705bf · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.216681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:636d126f6ccc0c23ddf70bb1dc13d57cce23a95d7591af8e1ba0bf7885c8a539

Observation 9d1b3b68-0735-4663-a951-812655a81338 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.188794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:b935d3b7134907fab16d36afbac2caf054b73ac8c6bc287febf2cc0da638e5c2

Observation 850ea5ce-203f-4cd4-ac50-d4befa26a7f0 · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.192029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:eef967f11626a1a9e83e2e842af7253384085cbd8657cb776fa4dfed4d03995b

Observation b7c49bf2-115f-4904-84c1-e6acad251b41 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.161743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:74c3768981f2b5439d813575ddc65f654f958d20e14e9f98302fa5e6a920a64c

Observation 8a68d7a6-8893-4eb9-a143-b4b0bfbc9d1f · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams VideoChat: Chat-Centric Video Understanding

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.156986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:cba1f37648720219c0617d9efe7e45f3bc683d5626e4f2e7c9634920d1327acc

Observation 8f3c6715-e53b-4656-b04f-e4ef8b672e7d · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:497f685849455601305336ef190ada3254f28340d674a92776e18338ebad1b11

Observation 896845f4-3d43-4c4e-b595-bd8a8b081d21 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.183828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:538863a5430a439abd39645f46887231dfe71c3db700606c191ccd178c2ae4f5

Observation f5d290d5-32f7-4035-aa05-f1d6cad28be8 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.159372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:484e3956dea8c0d78fcedae07ba5f1e19d4bbe7af09db2e63f56758c201d66b5

Observation 89a1a066-f3e4-47fc-9d6e-9f3370537c85 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.170346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:fd1e218636e88d75eceb3e21ba9ae7cfe6affbebb80baea4c92376801e2169d9

Observation b81ab108-434d-43d4-8619-da982bb9e121 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Timechat: A time-sensitive multimodal large language model for long video understanding,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:50d53726b0dce74a221cbc86ea58d34969067451502a63f699e1e0255aaec968

Observation 8b069b59-ddbf-4ffd-8919-b3520863a203 · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.143340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:62723b041f40827807d4c79530592bcb54f3b0db3e35e1c74f5eda8febfd9e19

Observation b17d68cf-e15c-4649-9cce-ea65ece727a3 · outbound

This paper cites MovieChat+: Question-aware Sparse Memory for Long Video Question Answering.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams MovieChat+: Question-aware Sparse Memory for Long Video Question Answering

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.118616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:005a1ca568c0c0a8bd2af38ff000241907b6fe46cad7eeafd1bc3161597528db

Observation 3beb55c8-9040-4568-9f47-17f7c3fbba7c · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:98ba5538c645244002a7d73de4632f701e17a80aed6b94f7944ba3c14ebdb03f

Observation 132ece48-ee35-4cf8-acfe-229f819960e3 · outbound

This paper cites Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:aaa826e5fceea927e90a70d0602698be7d1407c638423e88ea757f877d8d58c3

Observation 5650dc60-12d8-4aa7-91b8-e87c6fd5a55a · outbound

This paper cites Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:56.121280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:e5e9eb243d50c71f8b41911fe9a351255b1603b8d8708687f633a4b369bad241

Observation 9cc197e0-e82a-4cbf-9f40-d22dc731d015 · outbound

This paper cites Longvila: Scaling long-context visual language models for long videos,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Longvila: Scaling long-context visual language models for long videos,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:2c05dd0b455d87c2279f6d77eea2d68dff0270573385960cb60fc9a74b355a28

Observation b74f458d-aa20-4515-9c4f-45f2b761acae · outbound

This paper cites Long Context Transfer from Language to Vision.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Long Context Transfer from Language to Vision

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.151751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:bf0915a7df5a62abaf37b27d56447342aaf3da38f3576d23d044eae2dac78606

Observation e5176662-6b26-452f-b39e-1a184f94d697 · outbound

This paper cites Videollm-online: Online video large language model for streaming video,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Videollm-online: Online video large language model for streaming video,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:f4e53fd28c00cb480c74355c77a1b93ae7c11aa8316d03de8ea4f0ab16e5e45b

Observation 6481aeee-4d5d-4df6-992b-464ba936c4b9 · outbound

This paper cites Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:3ba09fc4584c640814434cf0c11bb2d8eaa460e14d329b24ce9e48b20b89504f

Observation 99fbe1a0-8d9e-4941-9c87-0bd4597ef252 · outbound

This paper cites LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.134846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:34ddaaa2862445ab513b09e8c4503c22ef7ddd560f4c1ab7e9556c9b3e116e3f

Observation 9787fc71-32b4-47b1-b609-6f718805aa53 · outbound

This paper cites StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.189303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:759ccb16be1b1f7953398d038aaea5a8d3ee197c555d1a359f076a1cbba9dcef

Observation b495186d-efd6-406a-829c-74dc87b044fb · outbound

This paper cites Videollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Videollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.465398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:94e67ffb6228febe7d25b2e4066a24a935ca2fd227e4b67c3315067c440d8a9f

Observation 2311b7f4-915f-4b60-8d5b-88630c2c589b · outbound

This paper cites Streaming long video understanding with large language models,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Streaming long video understanding with large language models,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:daf1fe3197783451dbfb9a863b958eadc5c4a75347f2a1181529c5198b5187a4

Observation 0f4abcfa-3cb7-4b83-b26d-38db5cfda108 · outbound

This paper cites Streaming dense video captioning,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Streaming dense video captioning,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:bb89346d885e5fcc7806bab4e20c2f2e3f87af26054ec811803fd875cee9a3f4

Observation a9e49ad1-317e-4670-b9c3-23633883a83e · outbound

This paper cites Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.211274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:738e9d69328ef7def1bda29117d97e61bcc70449fb01f06daffbef7e94db00b3

Observation 0500e553-c771-49b0-8cba-ee60a865a48a · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Ego4d: Around the world in 3,000 hours of egocentric video,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:b6fb6a9d65ce4e0917e9541c48a1fc9efe75154f1b8679a38bdc4ccf9b1b95c4

Observation 30980586-d4d5-4533-b4f2-57eb7e146384 · outbound

This paper cites Soccernet: A scalable dataset for action spotting in soccer videos,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Soccernet: A scalable dataset for action spotting in soccer videos,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:29b0de79edbc7267846dc1bc2395aa572c6a4ad6276098be306a91ad0c86aa4d

Observation 919e5318-5b27-451b-bf27-43c7421511b8 · outbound

This paper cites Svbench: A benchmark with temporal multi-turn dialogues for streaming video understanding.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Svbench: A benchmark with temporal multi-turn dialogues for streaming video understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.214126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:95d9b35a9403c24829befbd626d35ba08d96e7aaba1dde66ed30f5dfc01b033a

Observation 1f0a41bd-83de-4108-83c6-2735fe6e7321 · outbound

This paper cites OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.213701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:588f1d80234c7155dd8814ba08170462ece0e96f177367223a9f012f33798368

Observation 1bcb833d-577d-4561-af44-881459d45eae · outbound

This paper cites Livestar: Live streaming assistant for real-world online video understanding,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Livestar: Live streaming assistant for real-world online video understanding,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:863dd108f437feb53f4944d6a59d1e71a13db40c258603a7614eb24f7808135f

Observation 0c3c7d6c-8897-4c52-bbe5-8e1f27e12a82 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.196825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:8125bbcce411dd97ef1f8c99ebb7602831873bcba7ae3bffc7794713703a2e7c

Observation b06a6a4d-cdd9-4fad-9c35-2d1e87ea50c7 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Gemini: A Family of Highly Capable Multimodal Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.199389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:6d34130d575fbcdcb58e7c8f6cb6d0708c4d60730c88bbbcf681e7cfc8d8c1ec

Observation 9cb03962-87e3-4920-96a9-b01fc7faf7bc · outbound

This paper cites GPT-4 Technical Report.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams GPT-4 Technical Report

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.202622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:23b912e81f7ac329314ded076a97e750f22063d02313227b698175e044583e15

Observation 2e7b8826-60c1-4f8b-b441-0ab7bcc59934 · outbound

This paper cites Training language models to follow instructions with human feedback,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Training language models to follow instructions with human feedback,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:348e03dc7a393772afa519a96bd0e30abe610b4cc8fe5b4e48d601e0d683d737

Observation 05ff6f68-ea22-4834-bbf0-92723e921596 · outbound

This paper cites Improving language understanding by generative pre-training,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Improving language understanding by generative pre-training,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:32f7ba069f53e1bf84db1d338e59fabffa0509e432c7df8e9736d1827377a1da

Observation cd91ef08-29fb-4c34-9677-eb2dc798a317 · outbound

This paper cites Video dataflywheel: Resolving the impossible data trinity in video-language understanding,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Video dataflywheel: Resolving the impossible data trinity in video-language understanding,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:ad957cc71c772d93595f485b4c4f76ccb59f93a0e944b9b5af1f6c709abcc492

Observation e5671da8-4ac8-4e00-a6c0-3bd12e9e58d9 · outbound

This paper cites Object-centric rep- resentation learning for video scene understanding,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Object-centric rep- resentation learning for video scene understanding,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:e871f33faabdcd4e5a79b7b29d851c19ccd60e2aa4aaf81f4d311282e245609f

Observation 4491d186-d388-466b-859f-fabecf21408a · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Sharegpt4video: Improving video understanding and generation with better captions,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:d994d81006d94a8c7e9cc6689b01301896458ffebf70603c1b66bddbd224d2d8

Observation 1cc93271-321e-4114-8d60-8d17cb5e6254 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.205507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:034402306f4aa9f8153086738b83b044686b2a5d8ffd87eec30377bafcf580dc

Observation 9fecdd50-70a1-4ae9-ab49-82b7ad52c327 · outbound

This paper cites Video recap: Recursive captioning of hour-long videos,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Video recap: Recursive captioning of hour-long videos,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:0020ba5c8662e78e3e278a2d03c7b05994e05b4752fa1e322ddf4c0b96e06372

Observation 42234f78-d51b-40d0-9c58-0292d2ed1089 · outbound

This paper cites Large Language Models are Temporal and Causal Reasoners for Video Question Answering.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Large Language Models are Temporal and Causal Reasoners for Video Question Answering

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.460863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:4c5531f8d4c4cdc952fba0b4b9dfe44a990db320e16aaf132eb624f7f6a6f670

Observation 83444394-84c3-4d72-b9d5-00147384e322 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Mvbench: A comprehensive multi-modal video understanding benchmark,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:0b450f7cc22d6b736dbdbd138aa06b37055c2a283a73a25b18bd3ab24407db13

Observation 5a769220-b4bb-4308-8d91-65af6e1c98b5 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.200063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:51a468734b761a4b33e077aff3582262b7d310423e519ed40d1056e1cbaf3e78

Observation c968da58-0971-4b36-a17a-0caad4fe4b43 · outbound

This paper cites Learning to answer visual questions from web videos,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Learning to answer visual questions from web videos,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:a39817fb36993f176c6c32967ba918f2d960f7cbe0002deeb90a4e6b78892de1

Observation 636d20b2-a71e-4f72-b293-50aa4b974c99 · outbound

This paper cites Transformer-empowered invariant grounding for video question answering,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Transformer-empowered invariant grounding for video question answering,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:87fdcb74cfa0a6d900b56349f7b8b1d6626c374871e2efdec89aa73fde636373

Observation 23ca7f92-479d-4217-96f5-9a3837f06090 · outbound

This paper cites Intentqa: Intent question answering in videos by cognitive context reasoning,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Intentqa: Intent question answering in videos by cognitive context reasoning,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:424f3b1c9b98a08dc9933a33b4277d517c2c6013842e50942fc9430b167e764a

Observation eeba40fe-dce8-4d29-96f5-0b6a091ba3f3 · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:180cc587b567d55c83f537a25740b5d967a05c1666b21a19ffb5fbb9a7e382cf

Observation 1ee5b3c0-f9c4-4ee4-9339-814d4bd1a726 · outbound

This paper cites Vtg-gpt: Tuning-free zero- shot video temporal grounding with gpt,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Vtg-gpt: Tuning-free zero- shot video temporal grounding with gpt,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:7894dc46fb2debc14ee81d5ded8ce0051da4f609e4c7ba3c21d9223cd40a3584

Observation bc60d46a-72b6-452b-a97e-5da54f35f31d · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:56.197261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:c1c702e1b1c0c33414bfe04c208ad69a6ceacdfaee9921f00c5d8dd20d3de6ee

Observation 3e42a175-82f6-4606-bafb-acc2ecbbf98f · outbound

This paper cites Llava-next: A strong zero-shot video understanding model,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Llava-next: A strong zero-shot video understanding model,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:63e2f79221d2ab07a6a7c445bbf426ef50af944fe6edba37d3205161ef6de5e4

Observation 523f0f51-cc18-4bc2-a16a-7727ce90f3bd · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.194442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:391a0c0c975a47bbbe1915f2a31091341aa4b4e16d40aa8ea312fac1b28cc030

Observation f983d2b9-4280-42c8-bfe6-4f064420c107 · outbound

This paper cites Vila: On pre-training for visual language models,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Vila: On pre-training for visual language models,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:b659ef59cf7b9f3727689dd3ffbae05aa09f8e055fcca5697732ca33a8e7cb5c

Observation 39f2533c-8d43-4946-bb8c-3b45bfa48445 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.186287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:37278e84627438bdad28f49eb031f99d90218b16aefee7efb1ed2eef77bab42c

Observation 5ee2bf18-7555-468a-92a8-09edf6b7d236 · outbound

This paper cites Valor: Vision-audio-language omni-perception pretraining model and dataset,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Valor: Vision-audio-language omni-perception pretraining model and dataset,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:bddbad6ec5569a944ad0f024c2efb7d18c39cc12a2646261e89b84f6e1f8fe1e

Observation d9b9a6b8-d576-4173-a8bb-62cfc6fa0914 · outbound

This paper cites Cap4video++: Enhancing video understanding with auxiliary captions,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Cap4video++: Enhancing video understanding with auxiliary captions,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:3e80869f54ec94c40947d76471e3d53a4591c731b980125486293bb4fd22cf64

Observation 2a7516f1-14e1-4738-a662-9b7cb46fda77 · outbound

This paper cites Hierarchical banzhaf interaction for general video-language representation learning,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Hierarchical banzhaf interaction for general video-language representation learning,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:b8d1f213301b96ac6fc5fac22883e05cde595c2fb1413da6042f955cd6a80a0a

Observation 7710a38f-64b5-4172-9161-6f35811ed8d3 · outbound

This paper cites Video dataflywheel: Resolving the impossible data trinity in video-language understanding,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Video dataflywheel: Resolving the impossible data trinity in video-language understanding,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:b3938d9d665ce63fae42d9b7e8b48f6f83a832354056765658e3e6478c415294

Observation adb47ff5-f74e-4de9-87bb-1972c61e1e11 · outbound

This paper cites LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.194395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:8ebdfdd8fd49a6188fa9cd0982dd03910ecd80e5503e2d86534debaf6d097f34

Observation dd0c6591-e0ee-474c-b682-973138cd64dd · outbound

This paper cites Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.160538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:d281b451ce4468bcc5517d0d61bd8b0c9b63ee5afbf4da5f83ebed7e16bc2937

Observation 23438b1f-c851-45af-b843-b8d276f4c6de · outbound

This paper cites Querystream: Advancing streaming video understanding with query-aware pruning PREPRINT, 2026 18 and proactive response,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Querystream: Advancing streaming video understanding with query-aware pruning PREPRINT, 2026 18 and proactive response,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:450b23c5d16e317904842636d213f3e939c66caa180eb90b36917697c8b366ba

Observation 265f504f-45a7-4ff8-84fd-a238ec295f6d · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.172788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:bfbfd1d80fd3ac882ecb750d8083aa688eb91d640742bf644e354861828ceaf6

Observation a4363b46-8f26-4c53-9b55-1e9dbf016981 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Egoschema: A diagnostic benchmark for very long-form video language understanding,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:a973f4ff6b0f7282a7651510d21af1dac1b64ee584b727f8b28d8f3f98f0cd23

Observation 279110ff-7009-4def-9fc2-ef645bada4bf · outbound

This paper cites Activitynet- qa: A dataset for understanding complex web videos via question answering,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Activitynet- qa: A dataset for understanding complex web videos via question answering,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:97bcd4eb862caed8cee99c38e7b26e9aad0cbd982a436d5a8279a1efdb9cc276

Observation fe6319d5-aa72-41db-8405-77fa52d251b9 · outbound

This paper cites HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.175406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:54ec0d7f2ab4398b753b539d626d1b62d014f7c3ec2b596868aed357bcb00fdd

Observation a63644e3-1edd-4632-8e45-252edb43d375 · outbound

This paper cites Per- ception test: A diagnostic benchmark for multimodal video models,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Per- ception test: A diagnostic benchmark for multimodal video models,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:b4fcb0d0e4317f570171d2c351d0fcbbde2ec69bbacac773a35f40e0d2bb9de7

Observation a5450594-c4d0-4123-a469-43385212284b · outbound

This paper cites Social- iq: A question answering benchmark for artificial social intelligence,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Social- iq: A question answering benchmark for artificial social intelligence,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:2db9ed20eabd0f7327b4790688ad77cd1e60867a36e80ff9979a07461996f200

Observation e35b6132-bdf1-40c2-9e30-b80a8b5ec0b7 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Video question answering via gradually refined attention over appearance and motion,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:6d1c2accbee1f166f38c97fde349fdf448bc3f0ca44341e5b54d7980022638dc

Observation a9628186-9cfb-4c4f-89c7-83640a9e721b · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams TVQA: Localized, Compositional Video Question Answering

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.179069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:c924eb40a7057037b49cdfeec637bd8ca4248fcdcc96815738e2151dd1cd6784

Observation e51cedfa-774b-463a-8ff6-7d83cff31736 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Next-qa: Next phase of question-answering to explaining temporal actions,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:a055f427e6b13f0c21058bf76a4c53c68ac5954964a6b2e503fbd239432ea29c

Observation b076ea72-423e-4376-9fac-1e88a2ad387b · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Moviechat: From dense token to sparse memory for long video understanding,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:588e9b7eef15315373206465e7317c67b42a9d20946c177e56378059a6af644c

Observation 2302a3cc-dbea-4b04-9bf2-0a57450b0d54 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams LVBench: An Extreme Long Video Understanding Benchmark

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.186478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:ddba0b7eeb3e1332921f25edc9c2912637a079e1b0554f8f7a4f33b0d12c2185

Observation 65b7985b-bcdd-44be-98b5-5707d76db13e · outbound

This paper cites Tgif-qa: Toward spatio- temporal reasoning in visual question answering,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Tgif-qa: Toward spatio- temporal reasoning in visual question answering,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:49ee3dfa06732e75748ab12c85449545b33f22cba78bcac6a35fb9b936f450b5

Observation c505d8cd-bd90-40a0-b523-fedca3bcb250 · outbound

This paper cites Moviechat+: Question-aware sparse memory for long video question answering,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Moviechat+: Question-aware sparse memory for long video question answering,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:7d08cf2cafa1a25fbf4170730c65991ecb107e206462192e32bfbc9695a20a99

Observation 6632422c-0ed1-43c9-a40b-d79d16fbdaa0 · outbound

This paper cites Momentor++: Advancing video large language models with fine-grained long video reasoning,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Momentor++: Advancing video large language models with fine-grained long video reasoning,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:8776853bc389b5ab5f0569bdae661b8a3b61b6369b8adead9994765f5139c532

Observation 061d204b-0f3d-4a81-8946-65a586cc7203 · outbound

This paper cites Selongvlm: Empowering long video language models with self-corrective clip selection,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Selongvlm: Empowering long video language models with self-corrective clip selection,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:6ade1294fef910c606f600970242506fa800c7246e7a0da4ad5149e6176ad60d

Observation 4b419074-bb84-415c-858d-293a87258cff · outbound

This paper cites Ego-r1: Agentic chain-of-tool-thought for ultra- long egocentric video reasoning,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Ego-r1: Agentic chain-of-tool-thought for ultra- long egocentric video reasoning,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:f2c0ac7330e520023d693351320fc44bb8adbb1b17b186c2374047f532d03329

Observation e5494989-9c1a-43e9-a108-5219e788fd33 · outbound

This paper cites Hier-egopack: Hierarchical egocentric video understanding with diverse task perspectives,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Hier-egopack: Hierarchical egocentric video understanding with diverse task perspectives,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:59e7e7efb8243ab62576154688cd92875589b4f7e7b4f73d7277541b68949240

Observation 3062e5b6-9dfd-4c90-876c-a06af031dc4d · outbound

This paper cites Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.149055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:972a834571bfdf7c6bf98c7e418d4d3eacb0c820cca4792256e176b6881a6f63

Observation 9957af5d-bb0f-4d25-a5f3-838265ce4a72 · outbound

This paper cites Lvos: A benchmark for large-scale long-term video object segmentation,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Lvos: A benchmark for large-scale long-term video object segmentation,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:c39426c9a103eae14bd8c98b58ebd750b2c636149a6b0cac602832afefb74b2a

Observation fd2e873d-1e55-48ee-9dca-7ed7284483c2 · outbound

This paper cites Wild- video: Benchmarking lmms for understanding video-language interac- tion,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Wild- video: Benchmarking lmms for understanding video-language interac- tion,

Reference 83

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:1c0034789775ef6b14473c791b234a893987becca645ae636b1d73dce9cea956

Observation abf91a47-c61f-4aac-a7c3-05e74b05baf9 · outbound

This paper cites A survey on video temporal grounding with multimodal large language model,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams A survey on video temporal grounding with multimodal large language model,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:9e4fc74cd8fa9891a04435b7e1cde6c0dc309eb83b06090622a64086e402476f

Observation 3a690608-de23-42a1-b9d5-69b565545202 · outbound

This paper cites Timechat-online: 80% visual tokens are naturally redundant in streaming videos,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Timechat-online: 80% visual tokens are naturally redundant in streaming videos,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:0fb4981e227554c8e49f067ce875755647e8b7aa2333fb43b9743ca8a55195d1

Observation fd15ff1d-8407-4534-bb3b-6eae7f33d5d1 · outbound

This paper cites Videollamb: Long streaming video understanding with recurrent memory bridges,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Videollamb: Long streaming video understanding with recurrent memory bridges,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:061cc2b890ec32c58403741daaccc1e60c2aaa68fcd494ba3144cbb3bbe26791

Observation 0abc7edd-b519-4f5f-9ac9-331232039582 · outbound

This paper cites Evaluation by moments: Past and future,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Evaluation by moments: Past and future,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:cec1a9f9d2787e56494c46d500df49ce403b92ade2a0ccfe7480264ff1eb1c30

Observation d2ec4503-77ed-43ad-a76a-877f36af7609 · outbound

This paper cites Ldre: Llm-based diver- gent reasoning and ensemble for zero-shot composed image retrieval,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Ldre: Llm-based diver- gent reasoning and ensemble for zero-shot composed image retrieval,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:877652ed9c6763693d503dd23e411ecaee2745e4f715de730fa86adf8fdf0040

Observation 705236a4-f2cf-4dce-9ca9-6805022915cc · outbound

This paper cites Seman- tic editing increment benefits zero-shot composed image retrieval,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Seman- tic editing increment benefits zero-shot composed image retrieval,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:a3cc18891622a057a72ed9c62befd684acdc68e304b89516f5ed2d394200e29c

Observation 8513b66c-7f58-4cce-a5b5-b0fb7094c7be · outbound

This paper cites Visual instruction tuning,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Visual instruction tuning,

Reference 90

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:c839d42d9fa18286363405b2a5b54c546e8e89c075daea563d2753f91ae6f9bc

Observation f0186bb1-275d-4156-b913-c06e03da4361 · outbound

This paper cites Streamingcot: A dataset for temporal dynamics and multimodal chain- of-thought reasoning in streaming videoqa,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Streamingcot: A dataset for temporal dynamics and multimodal chain- of-thought reasoning in streaming videoqa,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:413c1e10bd470efba9eef3ed35462fc7f55e6bcfee554a0b67840abebb24493b

Observation cabb546d-af1a-46f3-8abe-4ff690e6b441 · outbound

This paper cites Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.151854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:068d96993540eef0b0a469dd194b96870cff17c813ecd16eda2e6910081c9efb

Observation 40caaa14-d8ec-4074-a426-ddab0dadf133 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.157736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:70e5ea961e33b1f108fb8c65fdffde5f4a4a356af841f767892d86b342999ecd

Observation bb3f9e80-9305-4707-9787-b669bfae4b17 · outbound

This paper cites InternLM2 Technical Report.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams InternLM2 Technical Report

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.181360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:4d2f0f15b933786c626a356f52f1a652c3ce0305684ef2cd0fcdf78bd7edad97

Observation 1138aca4-a08e-4f88-8c33-79fdb5d1b957 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.129188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:4c967b96720ab80ab952166811dbfa40e221e177af845d0961b78a5cfd496073

Observation 0f42cb28-47d0-458d-933f-f909572168a9 · outbound

This paper cites Qwen2.5-VL Technical Report.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Qwen2.5-VL Technical Report

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.137610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:ccf7d95ae55ac5e494d551c88a8d35e51705a6642705827d1ab0a8e045c85156

Observation 5ec331c7-0732-4199-b0e4-6caacad6680a · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 97

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:22ca68df5df06c2640a952880c2ddcef081218163b29d4e08cfabef793d022af

Observation 0cd7f72c-2c80-4c72-924f-9376be71d0e2 · outbound

This paper cites G-eval: Nlg evaluation using gpt-4 with better human alignment,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams G-eval: Nlg evaluation using gpt-4 with better human alignment,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:3f4a12b9efe2b04b25621b9819900bd3e6a8376fe374e874c625e96f8a649045

Observation 8bca9ac5-1dbd-4e0f-b11c-555f3302df39 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Longvideobench: A benchmark for long-context interleaved video-language understanding,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:44ce3365e55ab1fec9f66e684c89f073181d43ca81316e882e8afbbe705121bc

Observation 9d01f371-dd3b-4133-b7e5-095cbc80e3d0 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,

Reference 100

Resolution
unresolved
no resolver link, observed 2026-06-27T01:12:46.295455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:5385853ed2719722aa12bd26765e7322b626f089c0851e527c551ba20d94685d

Pith citing papers

No inbound Pith citation observations are available.