{"id":"af45873f-448c-478c-abc8-1403d3a18c94","arxiv_id":"2412.05725","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"BlackSwanSuite, a 15,469-question video benchmark for abductive and defeasible reasoning about unexpected events, shows the best VLMs underperform humans by 21-32%.","lead":"This paper introduces BlackSwanSuite, a video benchmark that tests AI models' ability to reason about unexpected events through abductive and defeasible questions. It finds that current vision-language models, including GPT-4o and Gemini 1.5 Pro, lag behind humans by up to 32% on these tasks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-annotator ground truth with no inter-annotator agreement threatens the validity of the reported model-human gaps.","rationale":"The reader's weakest assumption identifies both the single-annotator labels and the inflated human baseline. I focus on the former because it threatens the benchmark's construct validity: all MCQ/Y/N scores, the human baseline, and the hard-subset analysis are measured against labels generated by one person per video. The paper's own quality feedback admits invalid/valid confusions, yet no inter-annotator agreement is reported, so the uncertainty in the central 21-32% gaps is not quantified. A direct test is to use the two student annotators already collected in Appendix F.1: their per-question agreement is a free estimate of label reliability. If agreement is high, the concern is mitigated; if low, the benchmark needs consensus labeling before the gap magnitudes are taken at face value. This does not change the reader's CONDITIONAL verdict, which already requests inter-annotator agreement statistics; my analysis provides the specific test to satisfy it.","tokens_in":23044,"tokens_out":6738,"duration_ms":67788,"concrete_test":"Compute Cohen's kappa (or per-question agreement) between the two student annotators on the 100 MCQ and 150 Y/N questions they both answered in Appendix F.1, treating the original annotator's label as reference. If kappa < 0.7 or agreement < 80%, re-annotate a random 150-video subset with three independent annotators and recompute model accuracies against consensus labels; a material shift in the reported gaps would confirm that single-annotator labels are a load-bearing source of uncertainty.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, that models lag humans by up to 32% on abductive and defeasible video reasoning, depends on benchmark labels whose ground truth is produced by one annotator per video (Section 4.2). The quality check in Appendix B.2 covers only 60 of 1,655 videos, reports average correctness, depth, and grammar scores, and explicitly notes annotators sometimes 'marked a description that could have been valid as invalid (or vice versa).' No inter-annotator agreement statistic is reported anywhere. Because the Y/N variants label a single annotator's binary validation of their own hypothesis, and MCQ distractors are hypotheses that same annotator rejected, idiosyncratic or erroneous judgments are baked into the evaluation. Models that select a different but equally plausible explanation would be scored wrong, inflating the observed gap; a wrong validation could also misclassify the hard subset in Section 7.3. Without agreement data, the magnitude of the 21-32% gaps cannot be distinguished from annotation variance.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces BlackSwanSuite, a video reasoning benchmark built from 1,655 videos from the Oops! test set. Three tasks are defined — Forecaster (predict the future from the pre-event), Detective (abduce the hidden main event and validate hypotheses), and Reporter (describe the full video and validate hypotheses with full context) — with generative, multiple-choice, and yes/no variants totaling 15,469 questions. Six vision-language models (GPT-4o, Gemini 1.5 Pro, VideoChat2, VideoLLaMA 2, VILA-1.5, LLaVA-Video) are evaluated, and the headline finding is that the best models lag humans by 21–32 percentage points on MCQ and yes/no tasks, with additional analyses of perception/comprehension substitution, chain-of-thought prompting, and a hard subset. The data, validation/test splits, and leaderboard are publicly released.","tokens_in":23177,"tokens_out":6724,"duration_ms":64956,"significance":"If its measurement concerns are resolved, BlackSwanSuite fills a real gap: it targets abductive and defeasible reasoning about unexpected events, areas not covered by existing video QA benchmarks, and it provides a suite of tasks that decompose reasoning into prediction, abduction, and defeasible update. The paper is transparent about its data collection, reports a quality check, releases a leaderboard, and includes a useful diagnostic experiment (§7.1) showing that supplying human perceptual and comprehension annotations improves LLaVA-Video's MCQ accuracy. The potential value is high, but the core quantitative claim — the magnitude of model-human gaps — is not yet fully supported because the ground-truth labels rest on single-annotator judgments and the human baseline is constructed as a maximum over two annotators, as detailed in the major comments.","major_comments":[{"comment":"The ground-truth validity labels for the Y/N tasks and the correct/incorrect status of MCQ options come from a single annotator per video (Section 4.2), and the only reported quality check covers 60 of 1,655 videos with average correctness/depth/grammar ratings, not per-label agreement (Appendix B.2). Appendix B.2 itself records that annotators sometimes \"marked a description that could have been valid as invalid (or vice versa)\", and no inter-annotator agreement statistic is reported anywhere. Because a model that selects an alternative but equally plausible explanation is scored wrong, the reported 21–32% gaps cannot be distinguished from annotation ambiguity. Please report per-label agreement statistics (e.g., Cohen's kappa on validation decisions for a random sample of videos) and either re-annotate with multiple annotators or provide a multi-annotator subset on which the model-human gaps are recomputed.","section":"4.2, B.2, 7.3"},{"comment":"The human baseline is the maximum of two student annotators on 100–150 questions per task variant (Appendix F.1), which is an upper bound that inflates the headline gap; the main text (Section 5.2) states \"a human expert 150 questions for each task variant\" while Appendix F.1 says 100 for MCQ and 150 for Y/N, a direct inconsistency. The abstract's \"up to 32%\" gap in Reporter–Y/N is computed against this maximum baseline. Please report mean and median human accuracy over the two annotators alongside the max, provide confidence intervals for the sampled questions, and reconcile the stated number of baseline questions.","section":"F.1"},{"comment":"The hard subset in Section 7.3 is defined by the same single annotator's validity judgments — videos for which all Detective hypotheses were marked invalid in Reporter — so the hard/easy partition and the resulting 4.7–10.1 percentage point drops (Table 8) are not independent of the annotation noise documented in Appendix B.2. Please validate the partition with a second annotator or an alternative predictability measure (e.g., accuracy of human Forecasters on Vpre), or at least report the size of the hard subset and the overlap between annotators on the invalid-valid boundary.","section":"7.3"}],"minor_comments":[{"comment":"The sentence \"The Y/N variants for Forecaster (Detective) include each hypothesis proposed in step 1 (step 2)\" is confusing because Forecaster has no Y/N task according to Table 1; please rephrase to describe the two Y/N variants without the nested parentheticals.","section":"4.3"},{"comment":"The captions of Tables 2, 13, and 14 say \"Forecaster and Detective\" but the tables report Detective and Reporter results; please correct the captions.","section":"Tables 2, 13, 14"},{"comment":"The description of the human evaluators is inconsistent: Section 5.2 says \"a human expert\", while Appendix F.1 says two students; please clarify who the human benchmark participants were and how they were recruited.","section":"5.2, F.1"},{"comment":"No error bars or significance tests are reported for the model-human accuracy differences or for the perception/comprehension gains in Table 6 and the CoT gains in Table 7; given the small 100–150 item samples, please add confidence intervals or a significance test.","section":"6.1.1, 7.1, 7.2"},{"comment":"There are several typos and LaTeX artifacts: reference [2] contains \"Jastrzundefinedbski\", Section 8 has \"crucial step in models\" (missing article), and Figure 1 uses \"✅ explanation valid / ❌ explanation invalid\" without defining the symbols in the figure caption.","section":"References and text"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a serious benchmark contribution with substantial annotation effort and a public leaderboard, but the central claim depends on two measurement choices that need to be defended or revised: single-annotator validity labels and the max-of-two human baseline. I believe the authors can address these by adding agreement statistics and reporting mean human scores, so major revision rather than rejection is appropriate. The paper's fit to a computer vision venue is good; the domain of unexpected events is timely for VLM evaluation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: BlackSwanSuite is a worthwhile addition to video reasoning evaluation. It is the first benchmark I know that puts abductive and defeasible reasoning together on unexpected events, with three tasks and 15k questions built on a real video source. The construction is transparent and the evaluation is broad: six VLMs, closed and open, plus human baselines and some ablations. If you work on video QA or commonsense reasoning benchmarks, this is a dataset you should know about.\n\nCredit where due: the task design is thoughtful. Splitting videos into pre/main/post and asking Forecaster/Detective/Reporter tasks cleanly separates prediction, abduction, and revision. Building distractors from annotators' invalidated hypotheses is a neat idea, and the perception/comprehension ablation in §7.1 gives a concrete signal about where models fail. The appendix is honest about limitations, including potential training-data overlap with Oops! and the weaknesses of automatic metrics.\n\nSoft spots, in proportion. The load-bearing claim is the 21-32% gap over humans. I believe the gap is real in direction; every model is far below the human numbers across all four discriminative tasks. But the reported magnitude is inflated by two choices. First, the human baseline is max of two lab students, not average or majority, so it is an upper bound by construction. Second, ground truth comes from a single annotator per video, with no inter-annotator agreement reported anywhere; the 60-video quality check is decent but small, and the experts themselves noted that valid explanations were sometimes marked invalid. With Y/N labels being one annotator's binary call on their own hypothesis, idiosyncratic labels are baked in. Models that pick an equally plausible explanation get scored wrong. That does not destroy the benchmark, but it means the headline percentages should not be quoted as precise. The hard-subset analysis in §7.3 inherits the same issue.\n\nOne minor quibble: the LLM-Match metric uses Llama-3.1-8B to judge generations. It is a known evaluation choice and they report it, but it is weak for a paper whose main evidence is generative gaps; the human evaluation on 20 videos is too small to lean on. I would treat the generative results as suggestive, not conclusive.\n\nVerdict: send it to peer review. It deserves referee time. The referee should push for IAA statistics, a conventional human baseline, and ideally a released sample of multi-annotator labels. The benchmark is solid enough that these are revisions, not grounds for rejection.","headline":"A genuinely useful video reasoning benchmark whose headline human-model gaps are probably right in direction but softer than reported due to single-annotator labels and a max-of-two human baseline.","tokens_in":23727,"tokens_out":1898,"would_cite":true,"duration_ms":19143,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces BlackSwanSuite, a benchmark of 15,469 video-reasoning questions, and reports that state-of-the-art vision-language models lag human performance by up to 32 percentage points when reasoning about unexpected events.","keywords":["video reasoning benchmark","abductive reasoning","defeasible reasoning","vision-language models","unexpected events","commonsense reasoning","multiple-choice question answering","yes/no validation"],"falsifier":"Re-annotate a random sample of, say, 200 Detective–Y/N questions with five independent annotators each and measure agreement with the original single-annotator labels; if agreement is low or if a majority human vote moves the labels, the reported 32-point gap would need to be recomputed. A simpler check: run the human baseline with randomly chosen single annotators instead of the best of two experts; if the human score drops materially, the headline gap would be inflated.","tokens_in":22818,"feed_emoji":"🎬","tokens_out":7351,"duration_ms":65196,"temperature":0.7,"pith_summary":"BlackSwanSuite is a new benchmark for testing whether vision-language models can reason about unexpected events rather than recall patterns from training data. Each of 1,655 videos is split into before, main-event, and after segments, and the benchmark asks models to forecast what happens next, infer a hidden event from its surrounding context (abductive reasoning), and revise or confirm hypotheses when new video evidence appears (defeasible reasoning). Across 15,469 questions in multiple-choice, yes/no, and generative formats, the paper finds that the best evaluated models, including GPT-4o and Gemini 1.5 Pro, trail human experts by up to 32 percentage points on yes/no tasks and by about 25 points on abductive multiple-choice questions. The point of the benchmark is that atypical events are exactly where statistical recall fails, so the gap measures genuine reasoning limits rather than memorization.","feed_headline":"VLMs trail humans by 32% on reasoning about surprise videos","feed_subtitle":"A 15,469-question benchmark hides the key moment and asks models to infer it or revise their guess.","key_machinery":"The load-bearing mechanism is the three-part narrative decomposition of each video into $V_{\\text{pre}}$, $V_{\\text{main}}$, and $V_{\\text{post}}$, combined with a staged annotation process. Annotators first see only $V_{\\text{pre}}$ and propose what happens next; then they see $V_{\\text{post}}$ and mark those hypotheses valid or invalid while writing abductive explanations for the hidden $V_{\\text{main}}$; finally they see the full video and again validate or invalidate each hypothesis. Those validation labels become the ground truth for yes/no questions, the invalidated hypotheses become multiple-choice distractors, and the final explanations become the generative references. This structure turns defeasible reasoning into a concrete measurable operation: each hypothesis carries a timestamped defeasibility status that changes as visual evidence increases.","core_discovery":"The paper's central claim is that state-of-the-art VLMs cannot reliably perform abductive and defeasible reasoning about expectation-violating video events. To demonstrate this, the authors built BlackSwanSuite from 1,655 YouTube fail videos, each manually divided into pre-event, main event, and post-event parts, and annotated so that hypotheses proposed from earlier parts are later validated or invalidated once more video is shown. On the resulting tasks—Forecaster, Detective, and Reporter—the best closed-source models score 60.1–79.3% on multiple-choice and yes/no questions, while human experts score 85.3–95.3%, a gap of up to 31.9 points. The paper further shows that supplying human-written perception and comprehension descriptions improves one open-source model's Detective accuracy by up to 10 points, and that chain-of-thought prompting helps Detective but hurts Reporter, evidence that the bottleneck is distributed across perception, comprehension, and reasoning rather than a single component.","pith_inferences":["Editorial inference: because the reported human baseline is the maximum of two expert annotators, the published gap is an upper-bound estimate; a single average human may score lower, so the true gap could be smaller.","Editorial inference: the same three-part design could be applied to other event types, such as sports blunders, magic tricks, or animal behavior, to test whether the model-human gap is specific to the video source or reflects a general limitation in revising beliefs.","Editorial inference: the benchmark's yes/no defeasible questions resemble the belief-update steps needed in collaborative or assistive systems, so an immediate testable extension would be to measure whether models that pass here also revise a stated plan when given contradictory visual evidence.","Editorial inference: because some multiple-choice distractors were machine-generated from captions and edited by an LLM, part of the MCQ difficulty could come from stylistic confounds; re-running the best models on a fully human-authored subset would isolate reasoning from language-prior effects."],"forward_implications":["If the reported gaps hold, current production VLMs are not reliable enough for decisions that depend on revising an initial interpretation of a scene when new evidence appears, such as driving or surveillance.","The perception/comprehension ablation indicates that a large share of the failure is not in the language model but in grounding: giving models human-written descriptions of what is visible moves Detective accuracy by up to 10 points.","Chain-of-thought prompting is not a universal fix: it improves Detective multiple-choice accuracy but lowers Reporter accuracy, so its benefit depends on how much of the evidence is already visible.","On questions where even annotators could not guess the event until the video ended, model accuracy drops by up to 10.1 points, suggesting that the most surprising events are exactly the ones current models handle worst.","The generative evaluations show models often produce generic captions and miss the specific unexpected action, indicating the discriminative gaps reflect a broader understanding deficit, not just option selection."],"supporting_citations":[{"why":"Supplies the Oops! test-set videos with main-event localization that BlackSwanSuite filters and splits into pre/main/post parts.","marker":"[7]"},{"why":"Defines abductive commonsense reasoning and supplies the task framing that Detective builds on.","marker":"[3]"},{"why":"Provides the defeasible-reasoning logic (default reasoning) that motivates the validation and invalidation of hypotheses as new evidence appears.","marker":"[30]"},{"why":"The strongest closed-source baseline in the evaluation and the LLM used to edit machine-generated distractor captions.","marker":"[21]"},{"why":"A leading closed-source video VLM baseline; its 60.1% on Reporter Y/N anchors the 32-point gap to humans.","marker":"[33]"},{"why":"The strongest open-source baseline and the VLM used to generate pre-event captions that become machine-made distractors.","marker":"[43]"},{"why":"Open-source multi-frame VLM baseline used in the main evaluation.","marker":"[16]"},{"why":"Open-source video-LLM baseline evaluated across all task variants.","marker":"[4]"}],"fun_headline_variants":["VLMs flop on reasoning about surprise video events","Why AI can't reason about unexpected videos","BlackSwanSuite: AI fails to reason about surprise videos","VLMs need help with abductive reasoning in videos","Video AIs struggle to infer hidden events in unexpected clips"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ground truth labels for the 1,655 videos rest on a single annotator per video, with quality validation performed on only 60 videos, so the benchmark is only as reliable as that annotator's judgment of which hypotheses are valid.","fun_headline_variants_meta":{"raw":{"variants":["VLMs flop on reasoning about surprise video events","Why AI can't reason about unexpected videos","BlackSwanSuite: AI fails to reason about surprise videos","VLMs need help with abductive reasoning in videos","Video AIs struggle to infer hidden events in unexpected clips"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1349,"prompt_tokens":1028,"completion_tokens":321,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":644,"completion_tokens_details":{"reasoning_tokens":244}},"tokens_in":644,"tokens_out":321,"duration_ms":3601,"temperature":1.0,"reasoning_tokens":244,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:24:32.426937+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a random sample of, say, 200 Detective–Y/N questions with five independent annotators each and measure agreement with the original single-annotator labels; if agreement is low or if a majority human vote moves the labels, the reported 32-point gap would need to be recomputed. A simpler check: run the human baseline with randomly chosen single annotators instead of the best of two experts; if the human score drops materially, the headline gap would be inflated.","supporting_citations":[{"cited_title":"Oops! predicting unintentional action in video","cited_arxiv_id":null,"evidence_quote":"Supplies the Oops! test-set videos with main-event localization that BlackSwanSuite filters and splits into pre/main/post parts."},{"cited_title":"A logic for default reasoning","cited_arxiv_id":null,"evidence_quote":"Provides the defeasible-reasoning logic (default reasoning) that motivates the validation and invalidation of hypotheses as new evidence appears."},{"cited_title":"GPT-4o system card, 2024","cited_arxiv_id":null,"evidence_quote":"The strongest closed-source baseline in the evaluation and the LLM used to edit machine-generated distractor captions."},{"cited_title":"Video instruction tuning with synthetic data, 2024","cited_arxiv_id":null,"evidence_quote":"The strongest open-source baseline and the VLM used to generate pre-event captions that become machine-made distractors."},{"cited_title":"Vila: Efficient video-language alignment for video question answering","cited_arxiv_id":null,"evidence_quote":"Open-source multi-frame VLM baseline used in the main evaluation."}],"review_version":1}