{"paper":{"title":"Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL","cs.LG"],"primary_cat":"cs.CV","authors_text":"Aditya K Surikuchi, Alessandro Suglia, Andr\\'e F. T. Martins, Andr\\'e Viveiros, Ant\\'onio Farinhas, Baohao Liao, Beatriz Canaverde, Ben Peters, Chrysoula Zerva, Danae S\\'anchez Villegas, Desmond Elliott, Elena Bueno-Benito, Elias Stengel-Eskin, Emmanouil Zaranis, Giuseppe Attanasio, Jaehong Yoon, Mariella Dimiccoli, Miguel Moura Ramos, Mohit Bansal, Nithin Sivakumaran, Oswald Lanz, Pavlo Vasylenko, Raffaella Bernardi, Raquel Fern\\'andez, Sandro Pezzelle, Saul Santos, Shoubin Yu, Sonal Sannigrahi, Stella Frank, Vlad Niculae, Wafaa Mohammed","submitted_at":"2025-06-06T17:58:36Z","abstract_excerpt":"Despite recent progress in vision-language models (VLMs), holistic understanding of long-form video content remains a significant challenge, partly due to limitations in current benchmarks. Many focus on peripheral, ``needle-in-a-haystack'' details, encouraging context-insensitive retrieval over deep comprehension. Others rely on large-scale, semi-automatically generated questions (often produced by language models themselves) that are easier for models to answer but fail to reflect genuine understanding. In this paper, we introduce MF$^2$, a new benchmark for evaluating whether models can com"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2506.06275","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2506.06275/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}