{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:T7EF4SUOFTLLSQDQNLC2BWIURX","short_pith_number":"pith:T7EF4SUO","schema_version":"1.0","canonical_sha256":"9fc85e4a8e2cd6b940706ac5a0d9148de724540951e474f6b45f20ab24059c3a","source":{"kind":"arxiv","id":"2406.04264","version":3},"attestation_state":"computed","paper":{"title":"MLVU: Benchmarking Multi-task Long Video Understanding","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"MLVU benchmark shows current multimodal models struggle with most long video tasks and degrade sharply on longer clips.","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.CV","authors_text":"Boya Wu, Bo Zhang, Bo Zhao, Junjie Zhou, Minghao Qin, Shitao Xiao, Tiejun Huang, Xi Yang, Yan Shu, Yongping Xiong, Zheng Liu, Zhengyang Liang","submitted_at":"2024-06-06T17:09:32Z","abstract_excerpt":"The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially the insufficient lengths of videos, a lack of diversity in video types and evaluation tasks, and the inappropriateness for evaluating LVU performances. To address the above problems, we propose a new benchmark called MLVU (Multi-task Long Video Understanding Benchmark) for the comprehensive and in-depth evaluation of LVU. MLVU presents the following critical "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":true},"canonical_record":{"source":{"id":"2406.04264","kind":"arxiv","version":3},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","primary_cat":"cs.CV","submitted_at":"2024-06-06T17:09:32Z","cross_cats_sorted":["cs.AI","cs.CL"],"title_canon_sha256":"18a5bf14850be2ac26f321769b1dc07c5948b5228a893e1a555f8f75b88e6ccd","abstract_canon_sha256":"41148e2d78f4d31b4e397c1a844d4b6f5ce24671e5f8064a2655bd442342726a"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-05-17T23:39:21.958888Z","signature_b64":"KTb9S0ZdoT5X7XkVuO1YvtcJ6lef3HAFS3b7z6E2zvRhYDsraWyyuyAtIBuVa27fzEb/GiJfZqTZnpQUzBx6BA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"9fc85e4a8e2cd6b940706ac5a0d9148de724540951e474f6b45f20ab24059c3a","last_reissued_at":"2026-05-17T23:39:21.958186Z","signature_status":"signed_v1","first_computed_at":"2026-05-17T23:39:21.958186Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"MLVU: Benchmarking Multi-task Long Video Understanding","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"MLVU benchmark shows current multimodal models struggle with most long video tasks and degrade sharply on longer clips.","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.CV","authors_text":"Boya Wu, Bo Zhang, Bo Zhao, Junjie Zhou, Minghao Qin, Shitao Xiao, Tiejun Huang, Xi Yang, Yan Shu, Yongping Xiong, Zheng Liu, Zhengyang Liang","submitted_at":"2024-06-06T17:09:32Z","abstract_excerpt":"The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially the insufficient lengths of videos, a lack of diversity in video types and evaluation tasks, and the inappropriateness for evaluating LVU performances. To address the above problems, we propose a new benchmark called MLVU (Multi-task Long Video Understanding Benchmark) for the comprehensive and in-depth evaluation of LVU. MLVU presents the following critical "},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"The empirical study with 23 latest MLLMs reveals significant room for improvement in today's technique, as all existing methods struggle with most of the evaluation tasks and exhibit severe performance degradation when handling longer videos.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"The chosen video lengths, genres, and tasks in MLVU sufficiently represent the core challenges of real-world long video understanding and that performance on these tasks generalizes beyond the benchmark.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"MLVU is a new benchmark for long video understanding that uses extended videos across diverse genres and multi-task evaluations, revealing that current MLLMs struggle significantly and degrade sharply with longer durations.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"MLVU benchmark shows current multimodal models struggle with most long video tasks and degrade sharply on longer clips.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"c886ba5b9cf2cbdfb5ba0c78b921da150b529af8bb67e4b124e6a3446d2abb83"},"source":{"id":"2406.04264","kind":"arxiv","version":3},"verdict":{"id":"49f26cf4-8e6d-424a-beee-af7ca6b1887f","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-14T19:49:55.465952Z","strongest_claim":"The empirical study with 23 latest MLLMs reveals significant room for improvement in today's technique, as all existing methods struggle with most of the evaluation tasks and exhibit severe performance degradation when handling longer videos.","one_line_summary":"MLVU is a new benchmark for long video understanding that uses extended videos across diverse genres and multi-task evaluations, revealing that current MLLMs struggle significantly and degrade sharply with longer durations.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"The chosen video lengths, genres, and tasks in MLVU sufficiently represent the core challenges of real-world long video understanding and that performance on these tasks generalizes beyond the benchmark.","pith_extraction_headline":"MLVU benchmark shows current multimodal models struggle with most long video tasks and degrade sharply on longer clips."},"references":{"count":63,"sample":[{"doi":"","year":2023,"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","ref_index":1,"cited_arxiv_id":"2303.08774","is_internal_anchor":true},{"doi":"","year":2024,"title":"Anthropic. Claude 3. https://www.anthropic.com/ news/claude-3-family, 2024. 7, 2","work_id":"e9ea5ac3-366b-4128-9649-e6c18dfcadba","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2024,"title":"Minigpt4-video: Advancing multimodal llms for video understanding with interleaved visual-textual tokens","work_id":"cc937528-86d1-430f-bb5d-4980dbaadd72","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":null,"title":"Qwen Technical Report","work_id":"bb1fd52f-6b2f-437c-9516-37bdf6eb9be8","ref_index":4,"cited_arxiv_id":"2309.16609","is_internal_anchor":true},{"doi":"","year":2021,"title":"Frozen in time: A joint video and image encoder for end-to- end retrieval","work_id":"dbc0a634-c5d3-4207-8589-9084f58b919d","ref_index":5,"cited_arxiv_id":"","is_internal_anchor":false}],"resolved_work":63,"snapshot_sha256":"d9f5815a74603805b2696b4fe06a39032a108ead5b9be9d2cf5b4582c7eab1c0","internal_anchors":24},"formal_canon":{"evidence_count":2,"snapshot_sha256":"6625c7bfe019846c8f94968c384a39134d4c918e3f265a4782f40ab3d18a8b1d"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2406.04264","created_at":"2026-05-17T23:39:21.958307+00:00"},{"alias_kind":"arxiv_version","alias_value":"2406.04264v3","created_at":"2026-05-17T23:39:21.958307+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2406.04264","created_at":"2026-05-17T23:39:21.958307+00:00"},{"alias_kind":"pith_short_12","alias_value":"T7EF4SUOFTLL","created_at":"2026-05-18T12:33:37.589309+00:00"},{"alias_kind":"pith_short_16","alias_value":"T7EF4SUOFTLLSQDQ","created_at":"2026-05-18T12:33:37.589309+00:00"},{"alias_kind":"pith_short_8","alias_value":"T7EF4SUO","created_at":"2026-05-18T12:33:37.589309+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":64,"internal_anchor_count":64,"sample":[{"citing_arxiv_id":"2606.24187","citing_title":"Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling","ref_index":20,"is_internal_anchor":true},{"citing_arxiv_id":"2606.21734","citing_title":"HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning","ref_index":109,"is_internal_anchor":true},{"citing_arxiv_id":"2606.20726","citing_title":"How Well Can Your Video Model Remember? Measuring Memory-Budget Trade-offs in Long Video Understanding","ref_index":14,"is_internal_anchor":true},{"citing_arxiv_id":"2606.12195","citing_title":"InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning","ref_index":240,"is_internal_anchor":true},{"citing_arxiv_id":"2606.11740","citing_title":"UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA","ref_index":67,"is_internal_anchor":true},{"citing_arxiv_id":"2605.31529","citing_title":"SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence","ref_index":94,"is_internal_anchor":true},{"citing_arxiv_id":"2606.08615","citing_title":"Harnessing Streaming Video in the Wild","ref_index":73,"is_internal_anchor":true},{"citing_arxiv_id":"2606.05748","citing_title":"UNIVID: Unified Vision-Language Model for Video Moderation","ref_index":40,"is_internal_anchor":true},{"citing_arxiv_id":"2606.05917","citing_title":"MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering","ref_index":20,"is_internal_anchor":true},{"citing_arxiv_id":"2606.05833","citing_title":"Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models","ref_index":50,"is_internal_anchor":true},{"citing_arxiv_id":"2606.05769","citing_title":"Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction","ref_index":15,"is_internal_anchor":true},{"citing_arxiv_id":"2606.05259","citing_title":"VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding","ref_index":19,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03920","citing_title":"Benchmarking Visual State Tracking in Multimodal Video Understanding","ref_index":65,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03087","citing_title":"Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR","ref_index":146,"is_internal_anchor":true},{"citing_arxiv_id":"2606.07639","citing_title":"MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention","ref_index":41,"is_internal_anchor":true},{"citing_arxiv_id":"2605.14733","citing_title":"Video-Zero: Self-Evolution Video Understanding","ref_index":32,"is_internal_anchor":true},{"citing_arxiv_id":"2605.14906","citing_title":"MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models","ref_index":51,"is_internal_anchor":true},{"citing_arxiv_id":"2605.17921","citing_title":"An Efficient Streaming Video Understanding Framework with Agentic Control","ref_index":4,"is_internal_anchor":true},{"citing_arxiv_id":"2606.30220","citing_title":"From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA","ref_index":21,"is_internal_anchor":true},{"citing_arxiv_id":"2605.26680","citing_title":"DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding","ref_index":45,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24253","citing_title":"TuringViT: Making SOTA Vision Transformers Accessible to All","ref_index":78,"is_internal_anchor":true},{"citing_arxiv_id":"2605.30673","citing_title":"TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation","ref_index":25,"is_internal_anchor":true},{"citing_arxiv_id":"2605.31529","citing_title":"SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence","ref_index":94,"is_internal_anchor":true},{"citing_arxiv_id":"2412.04468","citing_title":"NVILA: Efficient Frontier Visual Language Models","ref_index":58,"is_internal_anchor":true},{"citing_arxiv_id":"2501.05067","citing_title":"LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding","ref_index":95,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":2,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/T7EF4SUOFTLLSQDQNLC2BWIURX","json":"https://pith.science/pith/T7EF4SUOFTLLSQDQNLC2BWIURX.json","graph_json":"https://pith.science/api/pith-number/T7EF4SUOFTLLSQDQNLC2BWIURX/graph.json","events_json":"https://pith.science/api/pith-number/T7EF4SUOFTLLSQDQNLC2BWIURX/events.json","paper":"https://pith.science/paper/T7EF4SUO"},"agent_actions":{"view_html":"https://pith.science/pith/T7EF4SUOFTLLSQDQNLC2BWIURX","download_json":"https://pith.science/pith/T7EF4SUOFTLLSQDQNLC2BWIURX.json","view_paper":"https://pith.science/paper/T7EF4SUO","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2406.04264&json=true","fetch_graph":"https://pith.science/api/pith-number/T7EF4SUOFTLLSQDQNLC2BWIURX/graph.json","fetch_events":"https://pith.science/api/pith-number/T7EF4SUOFTLLSQDQNLC2BWIURX/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/T7EF4SUOFTLLSQDQNLC2BWIURX/action/timestamp_anchor","attest_storage":"https://pith.science/pith/T7EF4SUOFTLLSQDQNLC2BWIURX/action/storage_attestation","attest_author":"https://pith.science/pith/T7EF4SUOFTLLSQDQNLC2BWIURX/action/author_attestation","sign_citation":"https://pith.science/pith/T7EF4SUOFTLLSQDQNLC2BWIURX/action/citation_signature","submit_replication":"https://pith.science/pith/T7EF4SUOFTLLSQDQNLC2BWIURX/action/replication_record"}},"created_at":"2026-05-17T23:39:21.958307+00:00","updated_at":"2026-05-17T23:39:21.958307+00:00"}