Pith. sign in

REVIEW 17 cited by

StreamChat: Chatting with Streaming Video

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.08646 v2 pith:2DQVLJV2 submitted 2024-12-11 cs.CV

classification cs.CV
keywords streamingvideointeractionstreamchatvisualcapabilitiescontentdecoding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents StreamChat, a novel approach that enhances the interaction capabilities of Large Multimodal Models (LMMs) with streaming video content. In streaming interaction scenarios, existing methods rely solely on visual information available at the moment a question is posed, resulting in significant delays as the model remains unaware of subsequent changes in the streaming video. StreamChat addresses this limitation by innovatively updating the visual context at each decoding step, ensuring that the model utilizes up-to-date video content throughout the decoding process. Additionally, we introduce a flexible and efficient crossattention-based architecture to process dynamic streaming inputs while maintaining inference efficiency for streaming interactions. Furthermore, we construct a new dense instruction dataset to facilitate the training of streaming interaction models, complemented by a parallel 3D-RoPE mechanism that encodes the relative temporal information of visual and text tokens. Experimental results demonstrate that StreamChat achieves competitive performance on established image and video benchmarks and exhibits superior capabilities in streaming interaction scenarios compared to state-of-the-art video LMM.

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Think in Sets for Streaming Video Token Compression

    cs.CV 2026-08 conditional novelty 7.0 of 10

    NovaCov uses a bounded, recency-weighted historical reference bank and a dual-branch submodular coverage objective to select streaming video tokens, outperforming training-free baselines.

  2. How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    Channel fusion gives better semantic grounding and QA performance in full-duplex LLM dialogue but is vulnerable to context corruption during interruptions, while cross-attention routing is more robust at the cost of w...

  3. Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?

    cs.CV 2025-11 unverdicted novelty 7.0 of 10

    Introduces the first dedicated benchmark for live multi-modal LLM task guidance with mistake detection and a streaming baseline model.

  4. Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

    cs.CV 2026-07 conditional novelty 6.5 of 10

    EgoMemo uses multi-scale temporal summaries, a knowledge graph, and visual archives to decide whether and when to intervene proactively on continuous egocentric video, setting baselines on the new EgoServe benchmark o...

  5. ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A training-free memory framework that anchors streaming video memory to latent objects discovered from frozen Video-LLM features, improving streaming QA accuracy while cutting memory and latency.

  6. ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Training-free latent-object memory anchors let frozen Video-LLMs retain object histories under a tight token budget and improve streaming and long-video QA.

  7. FOLIO: Focused Semantic Memory for Streaming Video Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Entity-centered focus-guided streaming memory lifts Qwen3-VL-8B to 82.0/69.1 Perception/Backward on OVO-Bench and 74.5 on StreamingBench while cutting writer tokens by ~32%.

  8. Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Stream3D-VLM adds autoregressive streaming control, VSFI geometry integration, GAVC compression, and a 1M-pair benchmark to enable real-time 3D VLM performance that beats prior models on 29 online and offline tasks.

  9. ProactiveLLM: Learning Active Interaction for Streaming Large Language Models

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    ProactiveLLM enables active interaction in streaming LLMs by learning semantic sufficiency cues from partial inputs through mask-based modeling and synchronized privileged self-distillation without external supervision.

  10. StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    StreamOV proposes evidence-guided long-short term memory and a hidden-state-driven trigger for efficient online audio-visual reasoning in streaming videos, along with the SOVBench benchmark for multi-turn evaluation.

  11. CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference

    cs.DC 2026-04 unverdicted novelty 6.0 of 10

    CodecSight reuses video codec signals for online patch pruning before the vision transformer and selective KV-cache refresh in the LLM, delivering up to 3x higher throughput and 87% lower GPU compute than prior baseli...

  12. ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    ViCoStream is a new coordinated pipeline framework for streaming VideoLLMs that achieves 134 FPS video throughput and less than 50 ms TTFT on A100 while keeping accuracy near full-history baselines.

  13. LiveVLN: Breaking the Stop-and-Go Loop in Vision-Language Navigation

    cs.RO 2026-04 unverdicted novelty 5.0 of 10

    LiveVLN enables smoother vision-language navigation by overlapping action execution with ongoing observation processing, preserving benchmark scores while cutting real-world waiting time by up to 77.7 percent.

  14. Existence of small semi-vortex solutions for the cubic nonlinear Schr\"{o}dinger system with Rashba type Spin-Orbit coupling on $\mathbb{R}^2$

    math.AP 2026-04 unverdicted novelty 5.0 of 10

    Small semi-vortex and ground-state solutions of the cubic NLS system with Rashba SOC on R² exist as energy minimizers under small mass, via concentration-compactness.

  15. Existence of small semi-vortex solutions for the cubic nonlinear Schr\"{o}dinger system with Rashba type Spin-Orbit coupling on $\mathbb{R}^2$

    math.AP 2026-04 unverdicted novelty 5.0 of 10

    Existence of small semi-vortex solutions for the Rashba SOC cubic NLS system on R^2 is proved via energy minimization under small mass constraint.

  16. cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations

    astro-ph.GA 2026-04 conditional novelty 5.0 of 10

    Recursive k-section domain decomposition, Morton-key hashing, and GPU dispatch cut communication and memory bottlenecks in RAMSES while preserving conservation to ~0.5%.

  17. cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations

    astro-ph.GA 2026-04 conditional novelty 5.0 of 10

    cuRAMSES replaces Hilbert-curve domain decomposition with recursive k-section partitioning and adds Morton-key hashing plus spatial binning to cut communication volume and accelerate feedback routines by up to 260x wh...

Pith tools