Pith. sign in

Paper Citation Record · LEDGER

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding

As of 8 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 1 inbound Pith citation observation for arXiv:2506.06275.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06275 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:00:57.046150Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T04:37:46.090777Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T06:06:24.603613Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact2
  • verified fuzzy26
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3dbbe518-562b-43e6-b03c-8de5cd7c6d21 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.837684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.837684Z digest=sha256:f2ab7673cb3892ea0b136d079779ff24508c153c2eb23456432c145315c6323c

Observation 95538f57-de00-4044-aec6-42e5ecb61d05 · outbound

This paper cites Infinibench: A comprehensive benchmark for large multimodal models in very long video understanding, 2024.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Infinibench: A comprehensive benchmark for large multimodal models in very long video understanding, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.841252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.841252Z digest=sha256:22de3ed232ff1bc66666343d78550c8a4b53dd2018de739947a3898b55addaeb

Observation 7fb0bc7b-3ca1-46ea-92d6-e076dd9f3974 · outbound

This paper cites Qwen2.5-VL Technical Report.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.843888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.843888Z digest=sha256:757e11a651873ad0654b5050c4744de573ff0d443af09ab3f060bdf90b0a1b90

Observation a3a76b66-8636-4821-b3f8-77c24e7b0085 · outbound

This paper cites Memory consolidation enables long-context video understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Memory consolidation enables long-context video understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.847080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.847080Z digest=sha256:bb182badfb605ea8b0149c89f3d3fa721548b437dd4739f2a7563d6a5845e6c8

Observation 179b2d20-9325-4ff1-b447-3b8fc0cc03b0 · outbound

This paper cites LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.849677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.849677Z digest=sha256:1923a172cc998bb40fc42fd7590a5df540ae0bd2b68cd78cd898d4708be017f6

Observation b2c254c6-ec9c-4f92-9412-132a96c5f277 · outbound

This paper cites Hadzic, Taran Kota, Jimming He, Cristobal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Fei-Fei Li.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Hadzic, Taran Kota, Jimming He, Cristobal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Fei-Fei Li

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.592958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.852611Z digest=sha256:3a57974bef3812234e19fa2503af2c1ef65afa732ff1e4b353ef366a7e12cead

Observation d426b64f-a31c-4848-bfb2-48061f7ec977 · outbound

This paper cites CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.855381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.855381Z digest=sha256:ddf2316cffbfe16315005f144f6f41030f4e3db86d5c0e782e2bc1e609d1eef7

Observation b4aafe70-e76b-4ea2-978e-0adce15a8fe8 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.858110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.858110Z digest=sha256:d08cd636abf94724108113dc642db84996aac25d83ae4b4dc6a6042d6ab3a64b

Observation b8268873-e78d-4ac4-b87e-286d9f98e358 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.860531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.860531Z digest=sha256:d67316785fc600a992e108a931f86e989461c5fd02dfc33aabe3654257028c8e

Observation e8a93478-79eb-43dc-9b5d-4b9eed8e11e0 · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.863411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.863411Z digest=sha256:75fa6d1189fa998e7c204dbb230fd9449d643bc5a72c092102dd3b4923a6b2a0

Observation 616ae0b2-6f27-4c8e-ad54-e29f750b57f8 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis, 2024.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.581669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.866203Z digest=sha256:aa3625cb07592fc7150d5f63702e821641a8a8ef9746582e8e61da8ffef08c37

Observation 88b87ae7-8177-4abd-ba3e-c8ff682564b9 · outbound

This paper cites Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.574381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.871356Z digest=sha256:f9e3ab2018133928e8364dbbc9bb2a1dbcfc8cfda392316e71152e1059dd3b6b

Observation 5ce38d0b-2dba-4664-bd79-57751542e4b5 · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Movienet: A holistic dataset for movie understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.566263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.876297Z digest=sha256:5413758b9127a2be914631d2a88d9e9fd351dccd6debb70ba541d69d847e0644

Observation a35c169a-c774-4c9d-98fa-faf3a21aa74c · outbound

This paper cites Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.557448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.878674Z digest=sha256:36d99c15052fa21a3e73b1ef7d42d61a4531824b3e439874992b7a05aba9dce2

Observation 23e8bf43-a19a-48ba-be62-8e24581ad37f · outbound

This paper cites Needle in a haystack - pressure testing LLMs, 2024.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Needle in a haystack - pressure testing LLMs, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.548694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.883819Z digest=sha256:30b813bcf5a5c81940ef60291867fb324cf517499cb0d2ee14958df3c3121311

Observation 43e4fb47-36f0-4e12-9846-706375296cd4 · outbound

This paper cites One thousand and one pairs: A “novel” challenge for long-context language models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding One thousand and one pairs: A “novel” challenge for long-context language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.886154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.886154Z digest=sha256:6704eceddb28eaf9958979552ab9b5d516bc1275696b0a504507d383f68f44fa

Observation 7f4ee9c5-c60f-4997-8830-9cd8754bbeba · outbound

This paper cites TVQA: Localized, compositional video question answering.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding TVQA: Localized, compositional video question answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.888635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.888635Z digest=sha256:bef9b0a3dfa84e737f6dd236240b9c2898f2d8b5e7ea9bd46b92464cbc86ab92

Observation 51d57200-a121-49dd-aee8-b6e92f6f5435 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.891037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.891037Z digest=sha256:31f3226b70c4d937482b65a9ebc5a869b97f44b2d959d7cc14c7a171dedf213e

Observation 75aff597-cc40-43a7-83aa-f1702c5a4ade · outbound

This paper cites Merlot reserve: Neural script knowledge through vision and language and sound.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Merlot reserve: Neural script knowledge through vision and language and sound

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.536206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.893414Z digest=sha256:93df8c8d098f7e117be0f1906318a8f14959aab51ca62b1a9a0ea6525b1e48e7

Observation 61a95390-c468-4b95-992e-c6f6636ede24 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.895831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.895831Z digest=sha256:3579bfc57e32a6297d4143a8130302860c1ce20e7ec7d323b99598b84f16b733

Observation d0161f92-a0c5-4a8c-a35f-243b89c77a80 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding VideoChat: Chat-Centric Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.898389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.898389Z digest=sha256:fbd6a02b552229497fd1ae2209a7a61734569c84659f068d0b3b880095871723

Observation 4eb44370-18c3-4b9d-9de5-3e902697d56d · outbound

This paper cites Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.901023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.901023Z digest=sha256:5f616bfd7a3c9412ab447b11cf73b44b45925105657c8e3357351ed9cf3bb43d

Observation 23218241-3d3b-4161-b4b2-224ebe739eb4 · outbound

This paper cites Contrastive decoding: Open-ended text generation as optimization.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Contrastive decoding: Open-ended text generation as optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.903725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.903725Z digest=sha256:893a00b9125bb1b976cfc6c7717a3ea92300ec1a322ed4943470d29816772d9f

Observation 4985de5e-2abd-4ad8-9b8b-15a208ca95b9 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models, 2023.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Llama-vid: An image is worth 2 tokens in large language models, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.906287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.906287Z digest=sha256:057f3d670adb195ff889bf89fa0eab164f306cd8822c781f573c0ea3a2aa8f4f

Observation 10def0f5-5245-457a-b658-355be0530ddf · outbound

This paper cites World model on million-length video and language with blockwise ringattention.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding World model on million-length video and language with blockwise ringattention

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.524542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.908532Z digest=sha256:793f4303687063fb7c152d3f5f9308a3099478c61f044afa9877985a116f1d71

Observation 967fb098-2b18-4567-a4d8-4143df8e4060 · outbound

This paper cites Visual instruction tuning, 2023.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Visual instruction tuning, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.910861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.910861Z digest=sha256:665eb689a1b5581d0abc61c7e936e8390f57b100001a8e9a765b27738046199e

Observation 3d98c38e-8bda-4777-b07f-04ba65877ee6 · outbound

This paper cites Is your video language model a reliable judge? InThe Thirteenth International Conference on Learning Representations, 2025.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Is your video language model a reliable judge? InThe Thirteenth International Conference on Learning Representations, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.513146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.913190Z digest=sha256:369c15c78ce2f9039c035b4321d5a51722507c72ce11ac9f3ab12ef59f2a4c1d

Observation 76eae7de-0c9e-4fe0-901d-087e9b3766f9 · outbound

This paper cites Nvila: Efficient frontier visual language models, 2024.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Nvila: Efficient frontier visual language models, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.505461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.915429Z digest=sha256:70fc107fba49580332c49bbd0eca6e30c89197cc72b665fd804436dbe30fbfa6

Observation dd4ac394-dfec-4cd4-8db4-7047bd0bcaf3 · outbound

This paper cites Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:00:57.241715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.917881Z digest=sha256:5e707b21737f89960711495a4cd1c729f10cb771ea35665a936645ad5f13d915

Observation e7ea8119-c001-4313-b8b4-2dd6293465b7 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.920387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.920387Z digest=sha256:cbf9d0494dca634092dadc4c7d4b47eae82334d91a1c4300640dfd451c515476

Observation 30c21deb-f754-42cd-8196-a1b5eeebe7c5 · outbound

This paper cites Valley: Video assistant with large language model enhanced ability, 2023.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Valley: Video assistant with large language model enhanced ability, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.497326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.923163Z digest=sha256:1ee5f5ffad52f933840280991cd248113eb15ec1ec04e20c0570a07a00dc7196

Observation 58b872c7-8679-4cdd-bba8-3cc549cdbb46 · outbound

This paper cites VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.925414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.925414Z digest=sha256:8f3ad83e9c78a492ec84d9aa5e0242e86a31e1e5738a15abf09914feb5e4da89

Observation bc071e67-7a13-4624-bd0d-15159caaee83 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.928000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.928000Z digest=sha256:dde60e92098e94b5ee5876945cd28ac445ab91f5f039788f858341c8388e7d5b

Observation d3ebb7c6-b8f5-4554-81cc-2595dd0a4c2b · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.930356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.930356Z digest=sha256:686ecd2f271aa368ec0dc8eafda92421c9f996da0b6f75dea2ce7ba55f4e9e87

Observation 98d7f9db-20c5-445b-8e17-87402094ef20 · outbound

This paper cites Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.933234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.933234Z digest=sha256:580b5d263bb32eb5e97abb20d8d32d0af6ddaf275e68de4cac768a296e12beed

Observation f37af5f2-89b8-4bcd-af6f-c0557e44a17c · outbound

This paper cites Neptune: The Long Orbit to Benchmarking Long Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Neptune: The Long Orbit to Benchmarking Long Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.935935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.935935Z digest=sha256:4d3fccf27784f92eb94c51e30b4f6b21bfd7e6ea92929a49556ac9839d121f87

Observation 65d1812c-b78e-4778-84de-5d824fb31caf · outbound

This paper cites GPT-4o System Card.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding GPT-4o System Card

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.938513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.938513Z digest=sha256:2e1d82d46e6e794d41f2644161c93e679ceb03ba795b1e700aa4597f6dc6693a

Observation 88b23c33-f7a5-4752-8eda-c3b0699c32fc · outbound

This paper cites Movie plot analysis via turning point identification.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Movie plot analysis via turning point identification

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.941552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.941552Z digest=sha256:fe0966862d374c334f6f882ad5853e5f2d9de44d670adaa0987fe680d1fef4c6

Observation b5976af9-30bf-423d-aa0e-455f78386ac3 · outbound

This paper cites Screenplay summariza- tion using latent narrative structure.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Screenplay summariza- tion using latent narrative structure

Reference 40

Resolution
verified exact
doi, observed 2026-08-07T06:00:57.084836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.944168Z digest=sha256:2b20089e6e6ae15e0bd0b02bbdcab45cb31c655e666c88fef42ffbdeaaff3320

Observation 5900ee82-62fd-4757-bae9-cfd7a231a0fc · outbound

This paper cites EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.946680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.946680Z digest=sha256:ce01a2c14dc0de88f21af2c7959ab95732ad7733c2c97093e9244b2dc515b151

Observation 886a64cd-0adb-4b7c-a86c-86f03b35f197 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Learning transferable visual models from natural language supervision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.949537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.949537Z digest=sha256:5fc321f028b2723ed76d3e2414dbc9d4e4568033df9b872520c0aae23984e38e

Observation 28df8268-ac92-432a-820f-81a3a296aedd · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Robust speech recognition via large-scale weak supervision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.951890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.951890Z digest=sha256:752d15842bd030d31ab7a6bc9d2701c6697f27549ddb9443260ab68c1caddde0

Observation 13e2c0fb-30ac-409e-9f52-7726c30ad6ae · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.954513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.954513Z digest=sha256:1fad19cb35315ea2ceddedc6c687203534f8bb85cfb758c3f060bb00fb2d4e25

Observation 6f61abf0-2a63-4875-8b84-cfeeb0d97834 · outbound

This paper cites $\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding $\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.957099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.957099Z digest=sha256:6a99e3cd178b4e0afc7c224e88549233bcfbbd2d935c91d3c2a63771b5723a9e

Observation ae2d1c1f-8cc0-49d6-834a-5a54be98c14a · outbound

This paper cites Trusting your evidence: Hallucinate less with context-aware decoding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Trusting your evidence: Hallucinate less with context-aware decoding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.479032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.959860Z digest=sha256:5d4ecc46eabdf51245f243943e88c9813b1ae56f35a5ee22509d2fce81b17996

Observation 835245b2-76a1-49c6-86be-dea589a25624 · outbound

This paper cites It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.964951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.964951Z digest=sha256:3fb0b0c89537aa58abf6bf9551c599c962ecf92270391283233ee9d80392f0ac

Observation 7e5fa39a-d49d-44f5-9ae8-4c8effec6e48 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.967687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.967687Z digest=sha256:323c4b855951cc2962e7dfd3384b7bf2cbc46689ca409b7d93790c380da4b432

Observation a72b4ca9-7698-4895-9ef0-3b6547706f43 · outbound

This paper cites MovieChat+: Question-aware Sparse Memory for Long Video Question Answering.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding MovieChat+: Question-aware Sparse Memory for Long Video Question Answering

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.970276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.970276Z digest=sha256:3cf4547aeea490ac312ea93a2221fac6789b0d62e6424989b21a6b31ec3cae48

Observation 075f24ad-0cd7-47cc-b7f8-d969bae98396 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.973042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.973042Z digest=sha256:d51b5ce84a0b62c2ab121dc229bfb47958ef6681b1bdf4711dfe496d1ef3efc2

Observation a695d49b-f432-4f8e-bcdb-a69c8560da44 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.975636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.975636Z digest=sha256:4e19bf27eb872116cb7f934c43783924291257e34f172740e947f6077b13b38d

Observation 08be5ddc-2573-4a46-9640-3a4072847048 · outbound

This paper cites AdaCAD: Adaptively decoding to balance conflicts between contextual and parametric knowledge.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding AdaCAD: Adaptively decoding to balance conflicts between contextual and parametric knowledge

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.471151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.978154Z digest=sha256:c1f8aba45b991715505f0ea6faa5ea8e7834fbc4af1696ba7707845fa0f315ff

Observation 7f3d84b0-1d89-45bf-afde-9d63081f84c0 · outbound

This paper cites Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.980537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.980537Z digest=sha256:6d8ab55bd4936beef665328c77e78c122fc763de224aa30a399eec548fee24e3

Observation 0bfbd83a-e957-481a-a63d-885e263d8b0d · outbound

This paper cites Lvbench: An extreme long video understanding benchmark, 2024.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Lvbench: An extreme long video understanding benchmark, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.983169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.983169Z digest=sha256:86c81806b95f5a34947ed7989cdc2048ba470b9b51e61985264a8eb7af44f2d7

Observation 9fb1d83c-2578-4f51-8ecc-0ca55ff4e8ab · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Videoagent: Long-form video understanding with large language model as agent

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.985963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.985963Z digest=sha256:b57b6049fc1747e3dcc7308e75707db1b7d04356e664a5ef4ce475aa9ad32f9f

Observation 462885d5-2bf7-4164-9e8f-e35caa3bab70 · outbound

This paper cites Videollamb: Long-context video understanding with recurrent memory bridges, 2024.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Videollamb: Long-context video understanding with recurrent memory bridges, 2024

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.455993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.988510Z digest=sha256:93fa85ab9cf26e4d1540b30bf2289e60cc56528227359c95b5d9ea85d63c53f0

Observation bafa5580-7d0f-462f-a507-2c1989dff989 · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.990818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.990818Z digest=sha256:f8b4f751a9c35b8ef94e52f040a2fbd2081036f981bbeb04766619e402f60432

Observation ccb18b42-d2ee-4817-86cf-18b7b39e0c50 · outbound

This paper cites Tenenbaum, and Chuang Gan.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Tenenbaum, and Chuang Gan

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.448759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:56.993378Z digest=sha256:8b88f564d26ef1c07a3b2744ebdaf4d0e5236f19ad744cac2db0e3686ee8cb8a

Observation 6fc27423-4722-45b9-9f75-2e30d3e446ed · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.995632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.995632Z digest=sha256:3de002657d1965e0579336ba56677fdd1cc8290198856d39f4997e92aee3c0fd

Observation 6760223f-58d2-4ec1-bbd7-148a5b92981f · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Next-qa: Next phase of question- answering to explaining temporal actions

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.997977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.997977Z digest=sha256:708531620695e0540b8926f0924b1d7fce98fda2e0d988796a5a0c136d23447e

Observation 0dbb8263-fde6-49b7-8c90-ba9d9797e36a · outbound

This paper cites Qwen2.5-Omni Technical Report.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Qwen2.5-Omni Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.000343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.000343Z digest=sha256:db25ea0c0ee6f50d949af7269ae666362b18a5350030ff1e0dd3d07be6c67b2b

Observation 64d2832e-edb9-4cf8-be2f-ef3b14498fc6 · outbound

This paper cites Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.002924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.002924Z digest=sha256:4cc93b323297d1585cf27a0585e51686bc5375ec9054dfae2c383fd097a352e9

Observation a62d46c6-1036-497c-96f4-38cdbb35b660 · outbound

This paper cites Just ask: Learning to answer questions from millions of narrated videos.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Just ask: Learning to answer questions from millions of narrated videos

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.430391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:57.005333Z digest=sha256:12bd4d325df0a9805cbdd9fc185ad0e06dc161197329d339a735c9d16d024eb4

Observation 9d7c74ba-3881-466f-8a23-3ed41b43d464 · outbound

This paper cites Justice or prejudice? quantifying biases in LLM-as-a-judge.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Justice or prejudice? quantifying biases in LLM-as-a-judge

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.423281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:57.007608Z digest=sha256:0baab15281e449b1ad88eb9e214f617d859e9595ff413380abfc0cf5ff9fc1f2

Observation c3633974-9125-46c9-90b8-b5523148329a · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.415781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:57.009993Z digest=sha256:77e7ea0b5367036cf8e55160ac816a8ffe564a2be4dd7d3b21ddf58da658a74d

Observation 50f3b593-8d14-4364-9b52-56505b331cd0 · outbound

This paper cites Merlot: Multimodal neural script knowledge models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Merlot: Multimodal neural script knowledge models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.408407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:57.012364Z digest=sha256:72825b22e0410b8294e0f17f5a36f743d62694b158661348aca2f851acc28971

Observation 1944a34a-1968-42dc-abc7-646a0bf944a0 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.014587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.014587Z digest=sha256:2b57a6195543e69cbbac687880f34f1790748a93eb5d84c00928b3932e40f9ff

Observation 10bd0a87-1047-4102-b7ed-356e09ae2a06 · outbound

This paper cites A simple llm framework for long-range video question-answering.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding A simple llm framework for long-range video question-answering

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.401439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:57.017356Z digest=sha256:10240588a5e3922de7de845078e6a317cbd1a5c1a559b598fa41b36ac42c7a76

Observation 1e4a762f-7acb-4b53-9d18-3bfc005b705b · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.019680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.019680Z digest=sha256:653c7aa9c4aa7e226db2ffd67f468796c4243a885d9ebcb43167eac7e4ab804a

Observation 21a9f31e-4cd3-4917-a883-0863f50c32fb · outbound

This paper cites LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.022279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.022279Z digest=sha256:6caa0869c63b652d19cbfc4e4b51f3e418749fe17a3cbf5547da03bc04309a6d

Observation 3b014608-1e4e-4c3d-bae1-f2fe67a62cb5 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.024929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.024929Z digest=sha256:5668221d5ed735a80f6ca954f18244b82a74439b311313c946c586e8f966b466

Observation a42027df-1fc4-417e-b038-935d905e835b · outbound

This paper cites Needle in a video haystack: A scalable synthetic evaluator for video MLLMs.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Needle in a video haystack: A scalable synthetic evaluator for video MLLMs

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.393887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:57.027321Z digest=sha256:0fad1e7591648a6ec3d332b984a416e4d2d070dfed3e173acdc41d6d6dde508f

Observation 1bede31b-70d3-40c3-9001-27aa1b135876 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.029729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.029729Z digest=sha256:513cb563b00807efcc265c43925afbe0e4f78dc75637137fee02bcdf6043fd85

Observation ef4cd2f1-7031-42bb-99f4-518c23f15f8a · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.032403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.032403Z digest=sha256:f8d83b66a8973464ce2b1d1e711bdb47b3d7b6bf2fe28a71e1d524120882fa68

Observation b500fd35-3dca-4ee8-9c01-3246711c998a · outbound

This paper cites The two claims should differ by minimal edits, meaning they should be as similar as possible while maintaining contrast.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding The two claims should differ by minimal edits, meaning they should be as similar as possible while maintaining contrast

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.386454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:57.035514Z digest=sha256:7816b51aafe393ae5116e328d5d6ec07ffe67949d62aa46065111dc8f0436fb5

Observation 9ae08d2f-83f7-4ef2-b045-6e00a89a15c4 · outbound

This paper cites Examples for Reasoning Granularity.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Examples for Reasoning Granularity

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.378904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:57.037997Z digest=sha256:65af673c7ce34313989b75fc6698e2dec44edbba1e8580877c404854026b4254

Observation 18497632-1550-42c2-84ff-5226ef626da0 · outbound

This paper cites Other" and suggest a new category. Note:The categorization is based on both claims (fact and fib). Check the examples provided in the “Examples for Comprehension Dimensions.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Other" and suggest a new category. Note:The categorization is based on both claims (fact and fib). Check the examples provided in the “Examples for Comprehension Dimensions

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.371793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:57.040534Z digest=sha256:e0a96eecd12b92073c0a72c7163e21f66ba4bf8333a824cf5d3f5aa6ef0b9b96

Observation dda476c8-6b53-4b78-8c16-5a594fce1e60 · outbound

This paper cites Pay attention to details and context in the movie, as some claims may be subtle or require careful reasoning.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Pay attention to details and context in the movie, as some claims may be subtle or require careful reasoning

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.364350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:57.043281Z digest=sha256:da151edb668cd0683ba5eb30738f0d0be70ee5d5ae337d6699cf097f33fb90e4

Observation d9a24bfa-ba3a-4c3b-8cc8-b974e1502508 · outbound

This paper cites Start Classifying Claims.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Start Classifying Claims

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.356777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:00:57.046150Z digest=sha256:688d1aa50b9480dbeecd74d5fee195318517ac025e44a7e5c8c62064b20e9d3e

Observation e2eb4797-bc4b-4baa-8094-4ec82f17dc74 · outbound

This paper cites doi: 10.18653/v1/2023.emnlp-main.308.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding doi: 10.18653/v1/2023.emnlp-main.308

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.881019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.881019Z digest=sha256:69f0e24814704750101aa3827d5a256cff5cac7cfa80273edd232e11c710a4b2

Observation 8dda0679-3ce2-4ba9-9350-a42df4b54a9f · outbound

This paper cites doi: 10.18653/v1/2024.naacl-short.69.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding doi: 10.18653/v1/2024.naacl-short.69

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.962238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.962238Z digest=sha256:30c64fb4890d75bf4a5e13234f99d21e9e11888a1a121d397e388da2f516eadc

Observation 83866043-2fe4-421c-baad-5889f5fab71d · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.873858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.873858Z digest=sha256:f69eacd290c89dd9a695848e58bd9be94915fdb2edc74451a6c3587ffbb934b5

Pith citing papers

Observation 0247282f-46b0-4ec8-a7af-5cfc4b9fa650 · inbound

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding cites this paper.

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:24.607724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:37:46.090777Z digest=sha256:4222d442147378f7fbb6488d2ae1aa1fdc99cc715cf4022073341aafeebf283d