{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:JTNSLMYZF2I32TPRLFO7RVYX7Z","short_pith_number":"pith:JTNSLMYZ","schema_version":"1.0","canonical_sha256":"4cdb25b3192e91bd4df1595df8d717fe7e098fc4012e717e69d41a5e643d5bac","source":{"kind":"arxiv","id":"2406.14515","version":3},"attestation_state":"computed","paper":{"title":"MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.MM"],"primary_cat":"cs.CV","authors_text":"Dahua Lin, Haodong Duan, Kai Chen, Kangrui Mao, Xiangyu Zhao, Xinyu Fang, Yining Li","submitted_at":"2024-06-20T17:26:01Z","abstract_excerpt":"The advent of large vision-language models (LVLMs) has spurred research into their applications in multi-modal contexts, particularly in video understanding. Traditional VideoQA benchmarks, despite providing quantitative metrics, often fail to encompass the full spectrum of video content and inadequately assess models' temporal comprehension. To address these limitations, we introduce MMBench-Video, a quantitative benchmark designed to rigorously evaluate LVLMs' proficiency in video understanding. MMBench-Video incorporates lengthy videos from YouTube and employs free-form questions, mirroring"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2406.14515","kind":"arxiv","version":3},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2024-06-20T17:26:01Z","cross_cats_sorted":["cs.MM"],"title_canon_sha256":"04bce401a98b75b80a991244d303a67b0a094bd772919b58b4d98b51b6d2eac2","abstract_canon_sha256":"2a4691b4ee1eabc09b2ec33d4e9278cff1da1f8cec0890ef0fc3887f5fd2c496"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:28:32.681104Z","signature_b64":"XZJjKN/LSenMmqOAZtBfUdgTD2ir4F9aSnglqNX+PbnkC5mgA1GVrHFZ+jCbbp1KdR4QTRyUa2ahMIB1qWfQCg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"4cdb25b3192e91bd4df1595df8d717fe7e098fc4012e717e69d41a5e643d5bac","last_reissued_at":"2026-07-05T09:28:32.680590Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:28:32.680590Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.MM"],"primary_cat":"cs.CV","authors_text":"Dahua Lin, Haodong Duan, Kai Chen, Kangrui Mao, Xiangyu Zhao, Xinyu Fang, Yining Li","submitted_at":"2024-06-20T17:26:01Z","abstract_excerpt":"The advent of large vision-language models (LVLMs) has spurred research into their applications in multi-modal contexts, particularly in video understanding. Traditional VideoQA benchmarks, despite providing quantitative metrics, often fail to encompass the full spectrum of video content and inadequately assess models' temporal comprehension. To address these limitations, we introduce MMBench-Video, a quantitative benchmark designed to rigorously evaluate LVLMs' proficiency in video understanding. MMBench-Video incorporates lengthy videos from YouTube and employs free-form questions, mirroring"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2406.14515","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2406.14515/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2406.14515","created_at":"2026-07-05T09:28:32.680653+00:00"},{"alias_kind":"arxiv_version","alias_value":"2406.14515v3","created_at":"2026-07-05T09:28:32.680653+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2406.14515","created_at":"2026-07-05T09:28:32.680653+00:00"},{"alias_kind":"pith_short_12","alias_value":"JTNSLMYZF2I3","created_at":"2026-07-05T09:28:32.680653+00:00"},{"alias_kind":"pith_short_16","alias_value":"JTNSLMYZF2I32TPR","created_at":"2026-07-05T09:28:32.680653+00:00"},{"alias_kind":"pith_short_8","alias_value":"JTNSLMYZ","created_at":"2026-07-05T09:28:32.680653+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":12,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2605.30673","citing_title":"TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2412.17574","citing_title":"HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2502.13923","citing_title":"Qwen2.5-VL Technical Report","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2407.03320","citing_title":"InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2502.04326","citing_title":"WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2501.12386","citing_title":"InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2501.04001","citing_title":"Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2501.13826","citing_title":"Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2406.07476","citing_title":"VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2504.10479","citing_title":"InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2412.05271","citing_title":"Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling","ref_index":65,"is_internal_anchor":false},{"citing_arxiv_id":"2508.18265","citing_title":"InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency","ref_index":33,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/JTNSLMYZF2I32TPRLFO7RVYX7Z","json":"https://pith.science/pith/JTNSLMYZF2I32TPRLFO7RVYX7Z.json","graph_json":"https://pith.science/api/pith-number/JTNSLMYZF2I32TPRLFO7RVYX7Z/graph.json","events_json":"https://pith.science/api/pith-number/JTNSLMYZF2I32TPRLFO7RVYX7Z/events.json","paper":"https://pith.science/paper/JTNSLMYZ"},"agent_actions":{"view_html":"https://pith.science/pith/JTNSLMYZF2I32TPRLFO7RVYX7Z","download_json":"https://pith.science/pith/JTNSLMYZF2I32TPRLFO7RVYX7Z.json","view_paper":"https://pith.science/paper/JTNSLMYZ","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2406.14515&json=true","fetch_graph":"https://pith.science/api/pith-number/JTNSLMYZF2I32TPRLFO7RVYX7Z/graph.json","fetch_events":"https://pith.science/api/pith-number/JTNSLMYZF2I32TPRLFO7RVYX7Z/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/JTNSLMYZF2I32TPRLFO7RVYX7Z/action/timestamp_anchor","attest_storage":"https://pith.science/pith/JTNSLMYZF2I32TPRLFO7RVYX7Z/action/storage_attestation","attest_author":"https://pith.science/pith/JTNSLMYZF2I32TPRLFO7RVYX7Z/action/author_attestation","sign_citation":"https://pith.science/pith/JTNSLMYZF2I32TPRLFO7RVYX7Z/action/citation_signature","submit_replication":"https://pith.science/pith/JTNSLMYZF2I32TPRLFO7RVYX7Z/action/replication_record"}},"created_at":"2026-07-05T09:28:32.680653+00:00","updated_at":"2026-07-05T09:28:32.680653+00:00"}