{"work":{"id":"ec969c94-54ed-47a7-8dd3-a07cd132485e","openalex_id":null,"doi":null,"arxiv_id":"2507.05240","raw_key":null,"title":"StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling","authors":null,"authors_text":"M","year":2025,"venue":"cs.RO","abstract":"Vision-and-Language Navigation (VLN) in real-world settings requires agents to process continuous visual streams and generate actions with low latency grounded in language instructions. While Video-based Large Language Models (Video-LLMs) have driven recent progress, current VLN methods based on Video-LLM often face trade-offs among fine-grained visual understanding, long-term context modeling and computational efficiency. We introduce StreamVLN, a streaming VLN framework that employs a hybrid slow-fast context modeling strategy to support multi-modal reasoning over interleaved vision, language and action inputs. The fast-streaming dialogue context facilitates responsive action generation through a sliding-window of multi-turn dialogues, while the slow-updating memory context compresses historical visual states using a 3D-aware token pruning strategy. With this slow-fast design, StreamVLN achieves real-time dialogues through KV cache reuse, supporting long video streams with bounded context size and inference cost. Experiments on VLN-CE benchmarks show state-of-the-art performance with low latency, ensuring robustness and efficiency in real-world deployment. The project page is: https://streamvln.github.io/.","external_url":"https://arxiv.org/abs/2507.05240","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-11T02:27:48.997495+00:00","pith_arxiv_id":"2507.05240","created_at":"2026-05-10T05:51:10.513833+00:00","updated_at":"2026-07-11T02:27:48.997495+00:00","title_quality_ok":true,"display_title":"Streamvln: Streaming vision-and- language navigation via slowfast context modeling","render_title":"Streamvln: Streaming vision-and- language navigation via slowfast context modeling"},"hub":{"state":{"work_id":"ec969c94-54ed-47a7-8dd3-a07cd132485e","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":32,"external_cited_by_count":null,"distinct_field_count":3,"first_pith_cited_at":"2025-11-21T09:52:07+00:00","last_pith_cited_at":"2026-07-07T02:42:41+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T16:49:41.238057+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":5},{"context_role":"baseline","n":1}],"polarity_counts":[{"context_polarity":"background","n":5},{"context_polarity":"baseline","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}