Pith. sign in

Paper Citation Record · LEDGER

StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2411.03628.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.03628 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.614557Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9bd4898b-b214-4606-8419-3699604de887 · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.254787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.254787Z digest=sha256:21b1c781b563a23c1c4410bb73241658444278d64f27ac36e022cd4f81f277f9

Observation 6a4356eb-06cf-4207-91f3-aa561ec377ba · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.296558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:f08f4fc8411919bbd4a546dd63b2204bf03e36f5a4b54753b309d3ebefb55fb6

Observation 9e7d65b3-1e6a-4799-a3bf-69b740edbb67 · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.614557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.614557Z digest=sha256:f4fc7a3a2301b4ad35f846091307981b36a42b921a43cc8f3602362f87d06b5f

Observation b345a33e-efb2-405e-b171-865a7a5412c5 · inbound

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale cites this paper.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.974973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.974973Z digest=sha256:5d3b3fbd41f89cbfdab451145f25629bfba4d1bf2c1340f987283833472aeaaa

Observation 63b43313-834b-48d5-8502-d609831f2d14 · inbound

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos cites this paper.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.248339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.248339Z digest=sha256:988778d5d433b15b92c22e04026e5149af1035b7afec0b1dbfda300989fb1564

Observation a7739048-e734-4807-b2d6-c769a53a1e94 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.849502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:f27b59ca985bca9815d06f671a009e5825119a5ae36daf3148eebd5c720dc2a5

Observation d7a362db-5f54-4548-b73e-99ceb56abe63 · inbound

Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models cites this paper.

Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:44.373465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:44.373465Z digest=sha256:a58870df0d6926a4ecd18656429343f02d502bde0f6155d778074c9c6fd1bfe9

Observation 1e5b8677-dccf-4afb-8f70-de7ab238b893 · inbound

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval cites this paper.

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:31:40.776721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T14:26:59.015559Z digest=sha256:5a1dd641be253b78afb42d8de81a5934a861e81c125485c28685cc482c3e1cd1

Observation c3ddb6c7-7078-487e-aae5-25a4dc1dd532 · inbound

AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction cites this paper.

AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T18:51:52.872171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:51:52.872171Z digest=sha256:0569edaf6c80785de14ab833f0a47debd68a40e2f4eb1838428e9087f8719f2c

Observation 06bc4f0d-0227-4fc5-9358-d442d0b8ea74 · inbound

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs cites this paper.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.479989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.479989Z digest=sha256:702118cf7e50856e26fba46a8ec491baa6ecbe7e45f4583c14102d2785dd54b8

Observation a49e3718-f95e-4026-a5f5-0ac0de340ec5 · inbound

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models cites this paper.

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.015160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.015160Z digest=sha256:08efc3b212d55d3c3ebcf560b34ac643ff322dc384671bbb58d5f861c8c3fa11

Observation 663612ab-4506-4272-a290-c619216eb2f4 · inbound

Diffractive electroproduction of light vector particles: leading Fock-state contribution in the presence of significant higher Fock-state effects cites this paper.

Diffractive electroproduction of light vector particles: leading Fock-state contribution in the presence of significant higher Fock-state effects StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T05:25:57.420224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:25:57.420224Z digest=sha256:1da3f155d52f61a0cd112347bb1da103d4445d6ba0855213cad164c2f2d7a52d

Observation 9b5af3be-a259-4fad-ac21-759bdd354340 · inbound

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos cites this paper.

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:51:28.116389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T02:49:12.987772Z digest=sha256:8e33aa263e10ce532701732eaad512cf9ca9576dd07d5a9b61f989965084c6c2

Observation 43fbf30b-2bc3-4141-acc6-7f7e6f52eb0c · inbound

Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously cites this paper.

Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T18:23:40.145298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:23:40.145298Z digest=sha256:42bf26d238890371b394bb067240545fa6eba693d7a1ff66871b65ddcd271ca2

Observation 1805d1dd-7ca4-41dd-96db-c66796f2cf02 · inbound

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance cites this paper.

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T22:09:56.270516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:09:56.270516Z digest=sha256:b435d6e4bc64713806f26be1a5d1e3881679736c298970d32163c977c7936326

Observation fd5e00ea-18bb-4c81-85b7-c98f527b91a1 · inbound

Seed1.8 Model Card: Towards Generalized Real-World Agency cites this paper.

Seed1.8 Model Card: Towards Generalized Real-World Agency StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:45:14.391226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T07:44:02.827006Z digest=sha256:0627fc1e4afa5ef222efe37b35ba7cb59af0119a0040aea97cfc40686ca121f5

Observation 8141243b-fd9f-4cb1-a5b3-ca88cf3d80cb · inbound

VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models cites this paper.

VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:25:58.057011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:39:30.731688Z digest=sha256:6f8b75a8c0e4399e813c05229de71aa50e9c3446d31ffc384bcd0225c8a3eeb4

Observation 8fbd045e-016c-4574-bda8-3852d9a6758d · inbound

Online Reasoning Video Object Segmentation cites this paper.

Online Reasoning Video Object Segmentation StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:31:00.212771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T15:29:03.843441Z digest=sha256:9993b429f2e2a1e6816f784ec4441c9937041bba21b4c9ef9be79d9b9a1b8c8d

Observation 5b490176-a817-4edf-8c78-e747aedccbaf · inbound

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning cites this paper.

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.757004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T06:51:52.861981Z digest=sha256:903681b3b97d29df3049e3f7f78b62037457a7db879c3713817ed37d74d86baa

Observation ac1758e3-3689-4543-8d90-2b3341a1d9a9 · inbound

Existence of small semi-vortex solutions for the cubic nonlinear Schr\"{o}dinger system with Rashba type Spin-Orbit coupling on $\mathbb{R}^2$ cites this paper.

Existence of small semi-vortex solutions for the cubic nonlinear Schr\"{o}dinger system with Rashba type Spin-Orbit coupling on $\mathbb{R}^2$ StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T18:51:26.143516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:51:26.143516Z digest=sha256:ec2a1cafb924c3d7e5cfc326f4ea9da404002a05e3fa85959a64b5605da0305c

Observation e5c4b23d-8c17-4e26-9a5b-4aa2d8ba3853 · inbound

LiveVLN: Breaking the Stop-and-Go Loop in Vision-Language Navigation cites this paper.

LiveVLN: Breaking the Stop-and-Go Loop in Vision-Language Navigation StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:13.476794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T01:51:23.379734Z digest=sha256:3e58373195f209f28a400078e896b3091258b0ea5b426c5e080508c11a708315

Observation 38540981-0d22-4fb4-a52e-02067ffebab7 · inbound

Don't Pause! Every prediction matters in a streaming video cites this paper.

Don't Pause! Every prediction matters in a streaming video StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.841108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T04:32:01.379605Z digest=sha256:da2a16b5ab51cadb9ab9d5a5636cbee44b30e5f0bcdc18de853aef958eefa6b9

Observation eb3684ab-300c-44e6-a7d0-7b730bf49737 · inbound

Decouple and Cache: KV Cache Construction for Streaming Video Understanding cites this paper.

Decouple and Cache: KV Cache Construction for Streaming Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:31:03.575498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T14:47:54.917408Z digest=sha256:1b1aaf1daa91c531c6911d21f2c2c978297085ae36eafd1ee6dc4c8632a3f345

Observation 4f062cc4-d96a-49cd-93af-920f45b444f0 · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:20:56.399096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T02:30:55.939351Z digest=sha256:fc38208ef7bdf1ca5738b2ea9334b8f6e7e76df8abd01968711420829a5255c6

Observation 8500b5a6-a38e-4324-a067-4887a00535ad · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.718872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:2cdef869de9109f1f0568e07961c779c91fc88a489ac3be187e5e55d3cb408af

Observation 66471926-6853-4272-a1f4-7383aab5a4e5 · inbound

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding cites this paper.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.707856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:21259b87d8138d568dd4b89912c68d39ae3c71d6a52b680b4a1f39335f7224e1

Observation 1659b41e-c847-4797-be5b-e7baf3a8462d · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.141874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T13:36:44.071188Z digest=sha256:c5734de8075fa19595718df0000372f42c20dd81462bc083ed188d381556db97

Observation 034d2306-1167-4490-a1f0-b510d06e5dd4 · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:19:20.357455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-04T01:11:42.073993Z digest=sha256:080531683367273f45a66aa6a6a9daaec4fd6e9048175f3b7a66bca3b1b71e4c

Observation 4da3be42-8592-4a21-83ac-9a5735db3fe0 · inbound

An Efficient Streaming Video Understanding Framework with Agentic Control cites this paper.

An Efficient Streaming Video Understanding Framework with Agentic Control StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.503209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T11:30:22.151045Z digest=sha256:684c648f22ea2f1fb016f96a1b90af78520ab6c711b697da40c5de9bb72a79dd

Observation ff030c80-ab6b-4b45-ba45-c3975a3dddef · inbound

MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering cites this paper.

MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:21:13.153199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T07:16:33.817183Z digest=sha256:46642af1d73ac70fa7a52c9ed061a229f196459e269fe7c6c04f5d6375e0fc2f

Observation ccc60d2c-bc99-4704-94b3-9b9082c406ef · inbound

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering cites this paper.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.070102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:8b96682393ba724c1e7f42e319a2a610b73ace4909026ce5aff20bc09b909794

Observation 154d3637-6094-4768-bc82-77fef37499fa · inbound

EGOSTREAM: A Diagnostic Benchmark for Streaming Episodic Memory in Egocentric Vision cites this paper.

EGOSTREAM: A Diagnostic Benchmark for Streaming Episodic Memory in Egocentric Vision StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.378541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T22:40:13.720341Z digest=sha256:63f628cfb3ab6a95342162ea9d8dbe166e62f2e605f785ffae45d1be49ad6fc5

Observation 9311291d-1d64-49b3-b8b7-1fe8401b5279 · inbound

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding cites this paper.

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:46:18.894558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T15:06:22.102725Z digest=sha256:3322d519fa37298484b4d09c836183ee90366b63406a3651bc15b99bf96d4a18

Observation c5407d40-6258-408b-b3bb-dd295f394749 · inbound

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding cites this paper.

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:44:36.402254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T10:42:37.401221Z digest=sha256:2f42d8bc477a28a2868b0555a3b956a0128a9e9c1e30d7b261ca1f57e53c656c

Observation b275d5f1-12b8-4bb1-a872-6cacc65ccf9d · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.871284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:3236743a65bf7acc462fa43822a0696bd174aaec0cf5a9f149f42bcbd6ac0870

Observation 447d702c-84d8-44c7-8173-805dae8e90c5 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.981195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:afc330f300f841c6cdc751a9f0044828f0c9b229f851635993116b5c0e345ab5

Observation 32ecbb8a-0066-4e7c-8ffa-0b790608bbff · inbound

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur? cites this paper.

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur? StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:30.642275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T16:52:22.811857Z digest=sha256:ab7bb8f98b42cfe9fecea368ec61f5e07fe7efb3d8e122ac3341741147d5c8c6

Observation 5ce92ae4-e5cd-491d-b047-f8cc06475ea6 · inbound

LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs cites this paper.

LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-26T07:29:08.001027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T07:26:07.356352Z digest=sha256:d0b423be5f7756d72a55c1de7cbc3f2715324d3e9c34bb11ca1d0083e71b9ce8

Observation ce35343c-1392-4842-8e43-433ed68de263 · inbound

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding cites this paper.

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:57.329945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T00:26:33.306128Z digest=sha256:5ffe2f511d34664be5e7d149fa289ec4784e06e0baba7d7924980f4785faf880

Observation 3f626472-c382-4986-a0cf-dcec7018dadd · inbound

ProtoKV: Streaming Video Understanding under Delayed Query with Summary-State Memory cites this paper.

ProtoKV: Streaming Video Understanding under Delayed Query with Summary-State Memory StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:19:50.783334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T05:18:38.341283Z digest=sha256:a85e9eade70ccd97abab4daf2fc9dc73d246e2fcf17f34fdda51a8359f62de71

Observation 4cd37fb4-85c1-4265-bc0c-a7e2aabd1774 · inbound

MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding cites this paper.

MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:38:39.538176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-03T16:37:09.666491Z digest=sha256:2230bddb38852607bdaead19c4e965ad901d818c29028f1abf920c898b278988

Observation 8d020d31-bed5-447b-a79b-05adf1fa0c91 · inbound

GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video cites this paper.

GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T05:36:40.278384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:36:40.278384Z digest=sha256:c9f6cedccd64767c6c0319e5fa338464f8108d4c49cdbed7b5538d0a0be8c9dd

Observation ab925452-f670-4ef3-848f-743a13b195e2 · inbound

QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding cites this paper.

QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T17:18:41.284513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:18:41.284513Z digest=sha256:cf9cb41f2ce93ef3476d928e1607c90bb6eb7c4983344e671986cc2ac847a499

Observation 95cf5144-b056-45b2-9ea3-86c4e8205783 · inbound

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos cites this paper.

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T05:01:06.200663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:01:06.200663Z digest=sha256:8684bb21f8b007958f87727bbf2b53d42a9c7163d6da1645c97c084d61971c60

Observation 9c3daa94-44a1-42f8-9aef-9bfe00ebec93 · inbound

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance cites this paper.

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T01:31:32.083457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:31:32.083457Z digest=sha256:d89ac4731fdf91db2d62b7bc50c2b6b3edb2f5db9f63fd0973ce82538ca5cc6f

Observation 6c5d6ee0-5824-4d04-8018-5c7fb7c2ab28 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.909328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.909328Z digest=sha256:0f82c7d545b287e8e4da51d4bab25285b5ea0c04901cf019091b7f253e49efbb

Observation dc85cc35-f2f5-44d9-9313-a0455b94904d · inbound

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding cites this paper.

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T03:21:44.907376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:21:44.907376Z digest=sha256:4e5bfe3d71a5eebb36461bf28bf46c3ce7238c55a1772b7e3b7f32f162c73a58

Observation 5dcb945f-5939-47d5-8b86-9410a0d91304 · inbound

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning cites this paper.

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T17:24:14.991460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:24:14.991460Z digest=sha256:88bf3ebce15055037c1ee9fbf2adf0d6ab1c8c828e492c67768f5969e081eb47