Pith. sign in

Paper Citation Record · LEDGER

StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2411.03628.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.03628 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.614557Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9bd4898b-b214-4606-8419-3699604de887 · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.254787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.254787Z digest=sha256:21b1c781b563a23c1c4410bb73241658444278d64f27ac36e022cd4f81f277f9

Observation 6a4356eb-06cf-4207-91f3-aa561ec377ba · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.296558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:6e07a61ce6306bf8fac85ed0cf9a31024db7e2bcc60b49c07df89d89d944b53c

Observation 9e7d65b3-1e6a-4799-a3bf-69b740edbb67 · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.614557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.614557Z digest=sha256:f4fc7a3a2301b4ad35f846091307981b36a42b921a43cc8f3602362f87d06b5f

Observation b345a33e-efb2-405e-b171-865a7a5412c5 · inbound

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale cites this paper.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.974973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.974973Z digest=sha256:5d3b3fbd41f89cbfdab451145f25629bfba4d1bf2c1340f987283833472aeaaa

Observation 63b43313-834b-48d5-8502-d609831f2d14 · inbound

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos cites this paper.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.248339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.248339Z digest=sha256:988778d5d433b15b92c22e04026e5149af1035b7afec0b1dbfda300989fb1564

Observation a7739048-e734-4807-b2d6-c769a53a1e94 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.849502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:3dadfca84fe1ab7f01ae4ebeba130ba3c267a230084d73f4894972ec905dd454

Observation d7a362db-5f54-4548-b73e-99ceb56abe63 · inbound

Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models cites this paper.

Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:44.373465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:44.373465Z digest=sha256:a58870df0d6926a4ecd18656429343f02d502bde0f6155d778074c9c6fd1bfe9

Observation 1e5b8677-dccf-4afb-8f70-de7ab238b893 · inbound

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval cites this paper.

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:31:40.776721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T14:26:59.015559Z digest=sha256:ecf4571e684863f5ea5f03e5d08a19238717a32088fe6f2a3ea3a392e885bdb8

Observation c3ddb6c7-7078-487e-aae5-25a4dc1dd532 · inbound

AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction cites this paper.

AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T18:51:52.872171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:51:52.872171Z digest=sha256:0569edaf6c80785de14ab833f0a47debd68a40e2f4eb1838428e9087f8719f2c

Observation 06bc4f0d-0227-4fc5-9358-d442d0b8ea74 · inbound

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs cites this paper.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.479989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.479989Z digest=sha256:702118cf7e50856e26fba46a8ec491baa6ecbe7e45f4583c14102d2785dd54b8

Observation a49e3718-f95e-4026-a5f5-0ac0de340ec5 · inbound

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models cites this paper.

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.015160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.015160Z digest=sha256:08efc3b212d55d3c3ebcf560b34ac643ff322dc384671bbb58d5f861c8c3fa11

Observation 663612ab-4506-4272-a290-c619216eb2f4 · inbound

Diffractive electroproduction of light vector particles: leading Fock-state contribution in the presence of significant higher Fock-state effects cites this paper.

Diffractive electroproduction of light vector particles: leading Fock-state contribution in the presence of significant higher Fock-state effects StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T05:25:57.420224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:25:57.420224Z digest=sha256:1da3f155d52f61a0cd112347bb1da103d4445d6ba0855213cad164c2f2d7a52d

Observation 9b5af3be-a259-4fad-ac21-759bdd354340 · inbound

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos cites this paper.

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:51:28.116389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T02:49:12.987772Z digest=sha256:57f963f4e89ff3259ad8e265e6147e0cf22907663dfc661326e23994038add8f

Observation 43fbf30b-2bc3-4141-acc6-7f7e6f52eb0c · inbound

Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously cites this paper.

Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T18:23:40.145298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:23:40.145298Z digest=sha256:42bf26d238890371b394bb067240545fa6eba693d7a1ff66871b65ddcd271ca2

Observation 1805d1dd-7ca4-41dd-96db-c66796f2cf02 · inbound

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance cites this paper.

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T22:09:56.270516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:09:56.270516Z digest=sha256:b435d6e4bc64713806f26be1a5d1e3881679736c298970d32163c977c7936326

Observation fd5e00ea-18bb-4c81-85b7-c98f527b91a1 · inbound

Seed1.8 Model Card: Towards Generalized Real-World Agency cites this paper.

Seed1.8 Model Card: Towards Generalized Real-World Agency StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:45:14.391226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T07:44:02.827006Z digest=sha256:c935f73448f8f83f6d843aafd994e327d8562c118009d630c1b86feb8057f1d6

Observation 8141243b-fd9f-4cb1-a5b3-ca88cf3d80cb · inbound

VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models cites this paper.

VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:25:58.057011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:39:30.731688Z digest=sha256:0d3f7ea305de131e0c55c6d9d1f782f20d27c84cdbf821607f8ba2f087b20d96

Observation 8fbd045e-016c-4574-bda8-3852d9a6758d · inbound

Online Reasoning Video Object Segmentation cites this paper.

Online Reasoning Video Object Segmentation StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:31:00.212771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:29:03.843441Z digest=sha256:05e8a51e0d83efdd5d7d881ea40193e7deac43c7a82a8a5cc9515fac05b93652

Observation 5b490176-a817-4edf-8c78-e747aedccbaf · inbound

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning cites this paper.

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.757004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T06:51:52.861981Z digest=sha256:9f9f477886a18c485d0d06d033407b557f16edea972552b997be2fac9f949df5

Observation ac1758e3-3689-4543-8d90-2b3341a1d9a9 · inbound

Existence of small semi-vortex solutions for the cubic nonlinear Schr\"{o}dinger system with Rashba type Spin-Orbit coupling on $\mathbb{R}^2$ cites this paper.

Existence of small semi-vortex solutions for the cubic nonlinear Schr\"{o}dinger system with Rashba type Spin-Orbit coupling on $\mathbb{R}^2$ StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T18:51:26.143516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:51:26.143516Z digest=sha256:ec2a1cafb924c3d7e5cfc326f4ea9da404002a05e3fa85959a64b5605da0305c

Observation e5c4b23d-8c17-4e26-9a5b-4aa2d8ba3853 · inbound

LiveVLN: Breaking the Stop-and-Go Loop in Vision-Language Navigation cites this paper.

LiveVLN: Breaking the Stop-and-Go Loop in Vision-Language Navigation StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:13.476794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T01:51:23.379734Z digest=sha256:bdb077b10524c6af25881f3f36927ad14ef5f7f081d09113eb768fccf8f3cfdf

Observation 38540981-0d22-4fb4-a52e-02067ffebab7 · inbound

Don't Pause! Every prediction matters in a streaming video cites this paper.

Don't Pause! Every prediction matters in a streaming video StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.841108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T04:32:01.379605Z digest=sha256:d18304d49f20e4aa99a61db6f3c8d409bd860b4ee58d06c2623ec033035ee954

Observation eb3684ab-300c-44e6-a7d0-7b730bf49737 · inbound

Decouple and Cache: KV Cache Construction for Streaming Video Understanding cites this paper.

Decouple and Cache: KV Cache Construction for Streaming Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:31:03.575498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T14:47:54.917408Z digest=sha256:02b162c5d30b2af7d20dfaabf0b06c256006b33b5e0089a8a31cba64e137f1e7

Observation 4f062cc4-d96a-49cd-93af-920f45b444f0 · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:20:56.399096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T02:30:55.939351Z digest=sha256:f28fede8d1bdd55df85b7f7b1b9e91b33496885e26b6fd7318fc65ec4b02df4a

Observation 8500b5a6-a38e-4324-a067-4887a00535ad · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.718872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:00538d4cd3dc62f48563d23ce1c52d3c5e75f7cd6bbd39673d55747563070847

Observation 66471926-6853-4272-a1f4-7383aab5a4e5 · inbound

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding cites this paper.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.707856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:ff8d2ab985adde25a28f2a1ae04e0569b212aac92208724063bd31e40396035b

Observation 1659b41e-c847-4797-be5b-e7baf3a8462d · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.141874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T13:36:44.071188Z digest=sha256:0fa561b55e0b3296d68719db7f349c093a02ca0ea5f279757ec5c19d353f462d

Observation 034d2306-1167-4490-a1f0-b510d06e5dd4 · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:19:20.357455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-04T01:11:42.073993Z digest=sha256:62868fa64ee33acd2c6e1f0c73b1a6d08d8cd1a65d3eba92c0dc015755cb0849

Observation 4da3be42-8592-4a21-83ac-9a5735db3fe0 · inbound

An Efficient Streaming Video Understanding Framework with Agentic Control cites this paper.

An Efficient Streaming Video Understanding Framework with Agentic Control StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.503209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T11:30:22.151045Z digest=sha256:bb7b0b0ec0bda8fbb28801d2d60b8a141e39c467968b280d81c0fbe4f45febca

Observation ff030c80-ab6b-4b45-ba45-c3975a3dddef · inbound

MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering cites this paper.

MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:21:13.153199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T07:16:33.817183Z digest=sha256:a49e09aee1216e03f8cb7259ebe4b0b58c4aad31187334f9e0f4774905099838

Observation ccc60d2c-bc99-4704-94b3-9b9082c406ef · inbound

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering cites this paper.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.070102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:03a3d86ed5cc0e60fe3015cf2b53f1b20ee6483a46ed0ebc9fdb18aea863688d

Observation 154d3637-6094-4768-bc82-77fef37499fa · inbound

EGOSTREAM: A Diagnostic Benchmark for Streaming Episodic Memory in Egocentric Vision cites this paper.

EGOSTREAM: A Diagnostic Benchmark for Streaming Episodic Memory in Egocentric Vision StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.378541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T22:40:13.720341Z digest=sha256:b87f8ae67e94e06b614f54f4838f097ea9e46ade52988f53bad7a4f97545a3ec

Observation 9311291d-1d64-49b3-b8b7-1fe8401b5279 · inbound

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding cites this paper.

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:46:18.894558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T15:06:22.102725Z digest=sha256:0e67ec2fe7f030309fa8233d0df8ac5adb824f10f38b39a70396a297e4801f92

Observation c5407d40-6258-408b-b3bb-dd295f394749 · inbound

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding cites this paper.

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:44:36.402254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T10:42:37.401221Z digest=sha256:2aa127f83bae2485dea47f50aaa01b270b3c9f24830c2534d75a0ea9654f90b6

Observation b275d5f1-12b8-4bb1-a872-6cacc65ccf9d · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.871284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:171b54fc10d94aa745c4cdc64f4d739796f03863eb6fd5943b877c0e7dbf4ebe

Observation 447d702c-84d8-44c7-8173-805dae8e90c5 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.981195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:174f63449adf25e543b3342036dc523644a380e6c7307fa6c21d57eb876c7bd3

Observation 32ecbb8a-0066-4e7c-8ffa-0b790608bbff · inbound

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur? cites this paper.

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur? StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:30.642275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T16:52:22.811857Z digest=sha256:410135ef625d8deae4854ff2632687ea1d47b7306c0e4bc17d887601f40ce26f

Observation 5ce92ae4-e5cd-491d-b047-f8cc06475ea6 · inbound

LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs cites this paper.

LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-26T07:29:08.001027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T07:26:07.356352Z digest=sha256:bbc4452dfc69af079aa19c41c8c9772670e9217a74edd9be3f83f52f3509700f

Observation ce35343c-1392-4842-8e43-433ed68de263 · inbound

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding cites this paper.

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:57.329945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:26:33.306128Z digest=sha256:29992e1c7bbb2742e34ecfa54372b713b072ab5853f1c178b53d319dd47d7590

Observation 3f626472-c382-4986-a0cf-dcec7018dadd · inbound

ProtoKV: Streaming Video Understanding under Delayed Query with Summary-State Memory cites this paper.

ProtoKV: Streaming Video Understanding under Delayed Query with Summary-State Memory StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:19:50.783334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T05:18:38.341283Z digest=sha256:d1647de856be6d5d396b0aa4ed4980c39421fab086456c7fe1f8a8bebfb1fe79

Observation 4cd37fb4-85c1-4265-bc0c-a7e2aabd1774 · inbound

MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding cites this paper.

MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:38:39.538176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-03T16:37:09.666491Z digest=sha256:a431b45de5a5619167e159bc47e21e5718c5c095541d87512d0b5fe766fa213a

Observation 8d020d31-bed5-447b-a79b-05adf1fa0c91 · inbound

GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video cites this paper.

GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T05:36:40.278384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:36:40.278384Z digest=sha256:c9f6cedccd64767c6c0319e5fa338464f8108d4c49cdbed7b5538d0a0be8c9dd

Observation ab925452-f670-4ef3-848f-743a13b195e2 · inbound

QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding cites this paper.

QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T17:18:41.284513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:18:41.284513Z digest=sha256:cf9cb41f2ce93ef3476d928e1607c90bb6eb7c4983344e671986cc2ac847a499

Observation 95cf5144-b056-45b2-9ea3-86c4e8205783 · inbound

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos cites this paper.

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T05:01:06.200663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:01:06.200663Z digest=sha256:8684bb21f8b007958f87727bbf2b53d42a9c7163d6da1645c97c084d61971c60

Observation 9c3daa94-44a1-42f8-9aef-9bfe00ebec93 · inbound

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance cites this paper.

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T01:31:32.083457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:31:32.083457Z digest=sha256:d89ac4731fdf91db2d62b7bc50c2b6b3edb2f5db9f63fd0973ce82538ca5cc6f

Observation 6c5d6ee0-5824-4d04-8018-5c7fb7c2ab28 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.909328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.909328Z digest=sha256:0f82c7d545b287e8e4da51d4bab25285b5ea0c04901cf019091b7f253e49efbb

Observation dc85cc35-f2f5-44d9-9313-a0455b94904d · inbound

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding cites this paper.

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T03:21:44.907376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:21:44.907376Z digest=sha256:4e5bfe3d71a5eebb36461bf28bf46c3ce7238c55a1772b7e3b7f32f162c73a58

Observation 5dcb945f-5939-47d5-8b86-9410a0d91304 · inbound

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning cites this paper.

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T17:24:14.991460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:24:14.991460Z digest=sha256:88bf3ebce15055037c1ee9fbf2adf0d6ab1c8c828e492c67768f5969e081eb47