{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:EZPHIONTL7C4EAWXGHT3EVP3OV","short_pith_number":"pith:EZPHIONT","schema_version":"1.0","canonical_sha256":"265e7439b35fc5c202d731e7b255fb755b050dbbea7d5cc17363a65740face3e","source":{"kind":"arxiv","id":"2412.10360","version":1},"attestation_state":"computed","paper":{"title":"Apollo: An Exploration of Video Understanding in Large Multimodal Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Felix Juefei-Xu, Licheng Yu, Nikhil Mehta, Ning Zhang, Orr Zohar, Philippe Hansen-Estruch, Serena Yeung-Levy, Tong Xiao, Xiaofang Wang, Xiaohan Wang, Xide Xia, Yann Dubois","submitted_at":"2024-12-13T18:53:24Z","abstract_excerpt":"Despite the rapid integration of video perception capabilities into Large Multimodal Models (LMMs), the underlying mechanisms driving their video understanding remain poorly understood. Consequently, many design decisions in this domain are made without proper justification or analysis. The high computational cost of training and evaluating such models, coupled with limited open research, hinders the development of video-LMMs. To address this, we present a comprehensive study that helps uncover what effectively drives video understanding in LMMs.\n  We begin by critically examining the primary "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2412.10360","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CV","submitted_at":"2024-12-13T18:53:24Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"f29365d62f90aee92569a0ec21d9f8b7bdda5160c8ed665b9a4a5fb193db1d5d","abstract_canon_sha256":"821047b9b08bce0833501f7266ff527ea7caeae7b685902999c56b7ab8ba7f32"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:48:56.865940Z","signature_b64":"i64qftNpREffE5i5OUav6emtdP0MVTGcltybllfra77FnSobl/IPebE7b8RbIbRm4yVV8WEovzUW1VXoGVGaDA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"265e7439b35fc5c202d731e7b255fb755b050dbbea7d5cc17363a65740face3e","last_reissued_at":"2026-07-05T09:48:56.865385Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:48:56.865385Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Apollo: An Exploration of Video Understanding in Large Multimodal Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Felix Juefei-Xu, Licheng Yu, Nikhil Mehta, Ning Zhang, Orr Zohar, Philippe Hansen-Estruch, Serena Yeung-Levy, Tong Xiao, Xiaofang Wang, Xiaohan Wang, Xide Xia, Yann Dubois","submitted_at":"2024-12-13T18:53:24Z","abstract_excerpt":"Despite the rapid integration of video perception capabilities into Large Multimodal Models (LMMs), the underlying mechanisms driving their video understanding remain poorly understood. Consequently, many design decisions in this domain are made without proper justification or analysis. The high computational cost of training and evaluating such models, coupled with limited open research, hinders the development of video-LMMs. To address this, we present a comprehensive study that helps uncover what effectively drives video understanding in LMMs.\n  We begin by critically examining the primary "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2412.10360","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2412.10360/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2412.10360","created_at":"2026-07-05T09:48:56.865440+00:00"},{"alias_kind":"arxiv_version","alias_value":"2412.10360v1","created_at":"2026-07-05T09:48:56.865440+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2412.10360","created_at":"2026-07-05T09:48:56.865440+00:00"},{"alias_kind":"pith_short_12","alias_value":"EZPHIONTL7C4","created_at":"2026-07-05T09:48:56.865440+00:00"},{"alias_kind":"pith_short_16","alias_value":"EZPHIONTL7C4EAWX","created_at":"2026-07-05T09:48:56.865440+00:00"},{"alias_kind":"pith_short_8","alias_value":"EZPHIONT","created_at":"2026-07-05T09:48:56.865440+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":9,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.12195","citing_title":"InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning","ref_index":263,"is_internal_anchor":false},{"citing_arxiv_id":"2605.31029","citing_title":"PEEK: Picking Essential frames via Efficient Knowledge distillation","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29472","citing_title":"Agent-Computer Observation Interfaces Enable Dynamic Computer Use","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2605.25979","citing_title":"LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence","ref_index":45,"is_internal_anchor":false},{"citing_arxiv_id":"2605.21625","citing_title":"Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly","ref_index":49,"is_internal_anchor":false},{"citing_arxiv_id":"2501.12386","citing_title":"InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling","ref_index":38,"is_internal_anchor":false},{"citing_arxiv_id":"2512.13511","citing_title":"Adapting MLLMs for Nuanced Video Retrieval","ref_index":92,"is_internal_anchor":false},{"citing_arxiv_id":"2501.13106","citing_title":"VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2604.17375","citing_title":"When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models","ref_index":79,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/EZPHIONTL7C4EAWXGHT3EVP3OV","json":"https://pith.science/pith/EZPHIONTL7C4EAWXGHT3EVP3OV.json","graph_json":"https://pith.science/api/pith-number/EZPHIONTL7C4EAWXGHT3EVP3OV/graph.json","events_json":"https://pith.science/api/pith-number/EZPHIONTL7C4EAWXGHT3EVP3OV/events.json","paper":"https://pith.science/paper/EZPHIONT"},"agent_actions":{"view_html":"https://pith.science/pith/EZPHIONTL7C4EAWXGHT3EVP3OV","download_json":"https://pith.science/pith/EZPHIONTL7C4EAWXGHT3EVP3OV.json","view_paper":"https://pith.science/paper/EZPHIONT","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2412.10360&json=true","fetch_graph":"https://pith.science/api/pith-number/EZPHIONTL7C4EAWXGHT3EVP3OV/graph.json","fetch_events":"https://pith.science/api/pith-number/EZPHIONTL7C4EAWXGHT3EVP3OV/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/EZPHIONTL7C4EAWXGHT3EVP3OV/action/timestamp_anchor","attest_storage":"https://pith.science/pith/EZPHIONTL7C4EAWXGHT3EVP3OV/action/storage_attestation","attest_author":"https://pith.science/pith/EZPHIONTL7C4EAWXGHT3EVP3OV/action/author_attestation","sign_citation":"https://pith.science/pith/EZPHIONTL7C4EAWXGHT3EVP3OV/action/citation_signature","submit_replication":"https://pith.science/pith/EZPHIONTL7C4EAWXGHT3EVP3OV/action/replication_record"}},"created_at":"2026-07-05T09:48:56.865440+00:00","updated_at":"2026-07-05T09:48:56.865440+00:00"}