Pith. sign in

Paper Citation Record · LEDGER

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering

As of 11 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2605.25621.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.25621 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:27:17.092553Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T18:17:53.013043Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T03:19:29.923486Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact20
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 169f237b-32aa-4abf-acba-d2b279e0472e · outbound

This paper cites Qwen2.5-Omni Technical Report.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Qwen2.5-Omni Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:02.128741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:e2f559b4066ed12ca7f5658ec6235dbb37b0c67b02f00660cbd3fd62a05edfe9

Observation c6ad0c3a-9735-4f59-938e-d003d854098a · outbound

This paper cites Qwen3-Omni Technical Report.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Qwen3-Omni Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:02.121593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:b768533742277deea830f9a09f4cc8187e1e43cf7e98cb9602cf8f4fdb8c75cc

Observation 40ab513a-9664-482d-980e-aab957d1e648 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Qwen2.5-vl technical report, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:4b95f1bb8f50d05927b9f66b4b3cfaee86d40b521e12ec2d6293e6f4955e6474

Observation 6fdc9f69-c0c2-4913-a6df-6f964c446a6a · outbound

This paper cites Qwen3-VL Technical Report.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Qwen3-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:02.120447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:6eef9559313b6993fe01666d8a3ef7346926f93f472bb4756d34e6efd0dbba57

Observation cd1981c4-98dc-4c92-bbfd-a2c421f95181 · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Videollm-online: Online video large language model for streaming video

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:724937e0ada6231b1dace019b9f7e874f99654e8c5f7f4664d14134b04b140af

Observation b532b3e5-3fe4-413f-9725-37d5115e37c6 · outbound

This paper cites Online video understanding: A comprehensive benchmark and memory-augmented method.arXiv e-prints, pages arXiv–2501, 2024.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Online video understanding: A comprehensive benchmark and memory-augmented method.arXiv e-prints, pages arXiv–2501, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:550a0aac3c12ea93488550006ae2e3cf0e05acc0576a08aaee6f7e0cec912f90

Observation 4a2f9743-d63c-4a06-8b9b-83b1964581a0 · outbound

This paper cites Streaming Video Question-Answering with In-context Video KV-Cache Retrieval.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Streaming Video Question-Answering with In-context Video KV-Cache Retrieval

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.124501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:4c59614724d86150be372f9017d5852c7215b7e1e05d5a45bebe2863df864ebe

Observation 1e70f0f2-cf34-40c3-a958-22be8f4e8bf0 · outbound

This paper cites Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.112392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:cf7966baf33d6ad7ded20a0486bc3721ada6655f74bad4cf55c4d598c7276f31

Observation fa9a1958-b772-49d3-a5f3-6ec2c5b6113c · outbound

This paper cites Streamforest: Efficient online video understanding with persistent event memory.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Streamforest: Efficient online video understanding with persistent event memory

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.117824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:2a18d2847782c6226a7fb5d64e76b4cb16988d71b1dc98753aa4c6440f1d2177

Observation 44eee74d-88f2-4cdd-abf1-72cc66110158 · outbound

This paper cites Streaming Video Instruction Tuning.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Streaming Video Instruction Tuning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:02.092580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:a89f9ed4934d656d0f70aa065b12659b650e2d0c85033bd9825a4fc02e45cb83

Observation 07960c7c-acc5-45c9-821e-b7b33dc352e9 · outbound

This paper cites Streammind: Unlocking full frame rate streaming video dialogue through event-gated cognition.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Streammind: Unlocking full frame rate streaming video dialogue through event-gated cognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:3584dd31e6c9ecb5af1de2d1cbd77d1dee1598628679f1e2f40f4059bb2d5041

Observation 3635b394-98ac-450c-a418-0428ec103da1 · outbound

This paper cites Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reaction.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reaction

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:568823fc5e9b37956a8b40f40a6d389627d8802e905384f727144de5eb6b1e35

Observation 73d34c97-5e41-49c6-96e1-f2106e8a44df · outbound

This paper cites Streamready: Learning what to answer and when in long streaming videos,.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Streamready: Learning what to answer and when in long streaming videos,

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.078593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:9ea70184a384a711abf9db209566abe3493fed802c28d97cad63f2f0f4585526

Observation ebb1124c-d1f7-4a5f-a6b9-c12848310fb1 · outbound

This paper cites Videochat: Chat-centric video understanding.Science China Information Sciences, 68(10):200102, 2025.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Videochat: Chat-centric video understanding.Science China Information Sciences, 68(10):200102, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:30ee4dc968e376fb5a98af6ffb48a333681d92f50acc97bf5aaaefadbbf05d4a

Observation c2145412-cd14-4103-a20e-0fe4fcf48399 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Video-llava: Learning united visual representation by alignment before projection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:1a7eb03f92f3205d3157e4fc110bcd03862981d64fdfea5f2c4532546aeee64b

Observation 8528f3dd-9cad-4af7-b93b-c19ef9be8b0a · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, April 2024.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Llava-next: A strong zero-shot video understanding model, April 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:c948d282393ac233cebd54dc61dae43e485b64f516774db6ec43105ce6333b2b

Observation 97775ab4-3dc8-4782-b1cb-b56c8a7f21c2 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:02.064876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:aeab5fe03c99fb7567b47762a64a9fbf84b247d585bb3460739b4feff2f4347b

Observation 8c758867-2559-4d80-a7be-ec27ec704beb · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:02.073382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:ec1e9f9559d5534fed5fcd6c2fb2e901265bfe37e66343bd42ac0679355cd8ce

Observation ea00e0e8-bebe-40a9-8fff-6dd28c05a162 · outbound

This paper cites Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation.Advances in Neural Information Processing Systems, 37:109922–109947, 2024.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation.Advances in Neural Information Processing Systems, 37:109922–109947, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:ccc1ac0893a72386553ba50420e19d907ee73785df6e9ce5e9ebb503ea58ebd2

Observation 54c6c3c8-1aea-418d-b4df-af03ab51a3ba · outbound

This paper cites Flash-vstream: Efficient real-time understanding for long video streams.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Flash-vstream: Efficient real-time understanding for long video streams

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:6f2387bdd1ab6a0da8f8f414d24bb391ccf5abbbbb35399bc50fe4674a9720d9

Observation 215ba10d-8f04-4e8e-8627-ef8e155eb9bd · outbound

This paper cites Timechat-online: 80% visual tokens are naturally redundant in streaming videos.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Timechat-online: 80% visual tokens are naturally redundant in streaming videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:c7962ff9a9375c078dcb8dcfa10789e13ceb44b8f191f81037f9abca4fe16ae0

Observation a8aa3ef4-d219-4cc9-a7c0-609d51473a4a · outbound

This paper cites StreamChat: Chatting with Streaming Video.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering StreamChat: Chatting with Streaming Video

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.061005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:8d91365ea50bb2b4141043c80872fabfca3cf420ca2dc8312db640139ab1d7ce

Observation ef958bc4-8acd-42f4-a146-d42ff1e7cbc1 · outbound

This paper cites LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:02.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:f7a95edc902bd1492f573a1b2689aadf2d536ca05492b4454b1880202984d856

Observation f970b990-1f10-43af-91a3-cf86e80520b0 · outbound

This paper cites Livecc: Learning video llm with streaming speech transcription at scale.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Livecc: Learning video llm with streaming speech transcription at scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:03aa23ddc70f90a51eaa4ee604276f48ea36944ee7af6e59da88fd7134d93105

Observation 14b4f96c-c2b2-4dda-9146-edf5da925d82 · outbound

This paper cites Roma: Real-time omni-multimodal assistant with interactive streaming under- standing,.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Roma: Real-time omni-multimodal assistant with interactive streaming under- standing,

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.060633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:88921956622c495e2cbe43f3bb50961902f98aca4ee0d6dae6e798926b7c7c0d

Observation 5b25ed02-b6a7-4dd2-890b-464239b3ef4b · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Moviechat: From dense token to sparse memory for long video understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:8c5dff7e3bbf7e1329ea63cecb3f9d6096dd53d624fb45ec63eb90abd67eb271

Observation c724b174-a450-42e7-82c2-db9d94cf9fed · outbound

This paper cites Ovo-bench: How far is your video-llms from real-world online video understanding? InProceedings of the Computer Vision and Pattern Recognition Conference, pages 18902–18913, 2025.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Ovo-bench: How far is your video-llms from real-world online video understanding? InProceedings of the Computer Vision and Pattern Recognition Conference, pages 18902–18913, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:ef9aededf6beae28e725aa1bb2b8f663bf06959610b6ae464dc66db3284f1a92

Observation ccc60d2c-bc99-4704-94b3-9b9082c406ef · outbound

This paper cites StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.070102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:835a04512faa59a61dfe74f0a2bd00737eb997ffae1832bdf5ac2dbd9eda67e4

Observation 28611be7-b521-4a79-b59f-77353832f72c · outbound

This paper cites Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.045071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:f1ebf5f7b05042ff298d29f917756b64eb2b82848279932f7c5db49303b0bf1a

Observation 945b7361-df98-4827-b89a-4dbdaeeb8869 · outbound

This paper cites Finevideo.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Finevideo

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:f20cc129d11e1830c0aa68bc7cfb279034bf809d036477031245b738ac004eb7

Observation d2575b81-030b-4026-bfa4-cc35bcc6833b · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:02.047325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:35185a67af9ecaef45a55c98612e3c416ba5dc3e63a6e769d3b6682f59de210a

Observation bc26f964-8476-4210-b91f-3110481eaaf7 · outbound

This paper cites Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-30T02:16:16.461780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:e60347b6f710d6f21fd713c817889de434d5155ab236900925ccd4c88f62f25e

Observation e41eb15f-f198-48ec-8970-35862b1c854f · outbound

This paper cites Learning transferable visual models from natural language supervision.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Learning transferable visual models from natural language supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:86fb1fe3c133294a0c6949f36d5fbd5ceb9a19e80293c086eefc904ad7b37431

Observation 1b8dcaf7-61a5-4e19-9c98-12c904288101 · outbound

This paper cites Clap learning audio concepts from natural language supervision.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Clap learning audio concepts from natural language supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:c4f0e256c254803932ebfbdc905aa571ffe9c34359949698982972640f5013da

Observation 8b6961c6-fb28-48f4-87a0-972e1ae5cf3c · outbound

This paper cites Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:02.050453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:f7106362dcbb9f56b821c8e8b88b7a934400c90ecee6dbd0dfd177ac811210b7

Observation 38b23bee-7f19-4a5a-9d59-0ec521497203 · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Mlvu: Benchmarking multi-task long video understanding

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:34:02.063984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:c783fe97907fcbf27df658c27a686e2f6756e7fed9d18899d23905900d62c2aa

Observation 05debda2-f479-45aa-9e00-15ccd78d2c0e · outbound

This paper cites Long Context Transfer from Language to Vision.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Long Context Transfer from Language to Vision

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:02.067305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:82b1dbec79f4d74756207fa45d14d15b68e72aa88f376a985a6e5b38e9148cfa

Observation 06199d9c-9c90-4ba0-bca6-6dfab8cdebb0 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.071338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:4c543e2c60aefc90bd6801576bd23303da38c61a821e00716d2449f244f27dd9

Observation f5115b0c-5cb2-4f1f-8278-f26d90da003c · outbound

This paper cites •0-2: static setup; talking head; minimal motion; no meaningful props.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering •0-2: static setup; talking head; minimal motion; no meaningful props

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:cddbf721600c0e3cff1e74fcef35b82d19905bbc238b1b81efb47d31a1f38d6c

Observation 8606ab11-5a12-4b9d-9f2a-c29761d939f1 · outbound

This paper cites •0-2: fragmented; contradictory; scenes don’t connect; ASR unintelligible.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering •0-2: fragmented; contradictory; scenes don’t connect; ASR unintelligible

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:1884f16c5f095ba06bd135d6f91fb9a897671ee9b406fb829d978b544a1bcd24

Observation 050c6f7e-1ecf-43d3-917b-ea00d0925d76 · outbound

This paper cites •0-2: mostly filler; few concrete entities; repetitive.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering •0-2: mostly filler; few concrete entities; repetitive

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:cd2848d9d997d3018987fe590916ee242190e5f1e0fb964df3658fdeeb8c4ce7

Observation f56556e7-284a-4721-8914-f08ec43eec89 · outbound

This paper cites •0-2: unrelated (generic voiceover vs visuals); frequent mismatches.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering •0-2: unrelated (generic voiceover vs visuals); frequent mismatches

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:6301b160dc7e92f7b6420a3bef6b9de42d2d27952c046e8fc6078721bcef9219

Observation 49d3bb98-7e57-4da3-846d-321cd6b9b0a7 · outbound

This paper cites video_uid.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering video_uid

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:22c46753f08c85b45a8cd723b49bf0ee24e9426303f96d5c88892cf6b06c0db4

Observation d5aac817-5dc7-4643-8d28-7ca760f85017 · outbound

This paper cites <Yes> " followed by the original assistant answer text. • Video path:.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering <Yes> " followed by the original assistant answer text. • Video path:

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:65d25f523d1de726829df3f635b47db340039e8399f9b03c15a66de0ecb2ffdd

Observation 389f07da-e2c4-4235-946a-7f76e1419e71 · outbound

This paper cites summary",.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering summary",

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-29T22:27:17.092553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:a2292b79217b208e0124c535658984bada369a4803cd6fd26336be91bd7e0bdc

Pith citing papers

Observation b97ac6e8-80aa-4a3e-9e93-87bf007ccd4b · inbound

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference cites this paper.

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:19:29.925891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T18:17:53.013043Z digest=sha256:b2b9c17261c18e28615da92ed351ca6d58f9339ea3e962415819b640159012dc