Pith. sign in

hub Baseline reference

World Model on Million-Length Video And Language With Blockwise RingAttention

Baseline reference. 50% of citing Pith papers use this work as a benchmark or comparison.

36 Pith papers citing it
Baseline 50% of classified citations
abstract

Enabling long-context understanding remains a key challenge in scaling existing sequence models -- a crucial component in developing generally intelligent models that can process and operate over long temporal horizons that potentially consist of millions of tokens. In this paper, we aim to address these challenges by providing a comprehensive exploration of the full development process for producing 1M context language models and video-language models, setting new benchmarks in language retrieval and new capabilities in long video understanding. We detail our long context data curation process, progressive context extension from 4K to 1M tokens, and present an efficient open-source implementation for scalable training on long sequences. Additionally, we open-source a family of 7B parameter models capable of processing long text documents and videos exceeding 1M tokens.

hub tools

citation-role summary

background 5 baseline 5 dataset 1 method 1

citation-polarity summary

representative citing papers

LVBench: An Extreme Long Video Understanding Benchmark

cs.CV · 2024-06-12 · accept · novelty 7.0

LVBench is a new benchmark for extreme long video understanding that evaluates multimodal large language models on hour-scale videos using tasks designed to probe extended memory and comprehension.

MLVU: Benchmarking Multi-task Long Video Understanding

cs.CV · 2024-06-06 · conditional · novelty 7.0

MLVU is a new benchmark for long video understanding that uses extended videos across diverse genres and multi-task evaluations, revealing that current MLLMs struggle significantly and degrade sharply with longer durations.

IOI: Decoupling Kinematics and Physics for Interactive World Models

cs.RO · 2026-06-22 · unverdicted · novelty 6.0

IOI decouples deterministic kinematics from stochastic physics in interactive world models by rendering forward-kinematics trajectories into multi-view projections that guide a video generator, achieving SOTA fidelity and OOD generalization on RoboTwin.

Lance: Unified Multimodal Modeling by Multi-Task Synergy

cs.CV · 2026-05-18 · unverdicted · novelty 6.0 · 2 refs

Lance presents a dual-stream mixture-of-experts model with modality-aware positional encoding and staged multi-task training that outperforms prior open-source unified models on image and video generation while keeping strong understanding performance.

citing papers explorer

Showing 36 of 36 citing papers.