{"id":"1ac0ca0d-466d-4394-8039-45aa57a830eb","arxiv_id":"2502.15683","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"VISIONE placed first in the 2024 Video Browser Showdown after a new question-answering task and a medical video dataset were introduced.","lead":"This short report lists the scores and rankings from the 13th Video Browser Showdown, a video retrieval competition held in Amsterdam in January 2024. It documents which of twelve interactive search systems performed best across expert and novice tasks, and introduces a new question-answering task type.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Manual judging of AVS/QAS submissions is the least verifiable step; without inter-judge consistency data, the score ranking and VISIONE's win rest on unmeasured judge variance.","rationale":"I read the paper as a transparent competition report, not a hypothesis-driven study. The central claim is a factual outcome: under the announced evaluation, VISIONE had the highest combined score. The raw data is archived, which is real independent support. The weakest point is the manual judging, because the paper itself identifies it as necessary for exactly the tasks that cannot be auto-checked, and because no reliability evidence is given. This is not a reason to reject the report—live judging is a legitimate part of the competition format—but it is a reason to treat the headline ranking as provisional until the archived judgments can be re-checked. The reader's weakest_assumption overlaps with this, though I place more weight on inter-judge consistency than on the DRES server, which is infrastructural and less likely to be biased. A targeted re-judging study would settle the question; absent that, the ACCEPT verdict stands.","tokens_in":7471,"tokens_out":10113,"duration_ms":92074,"concrete_test":"From the VBS-Archive raw export for VBS 2024, extract all manually judged AVS and QAS submissions. Have two independent judges re-score a stratified sample (e.g., 50 AVS and 50 QAS per session) against the original task definitions, blind to original verdicts. Compute Cohen's kappa and the resulting score differences. Then determine the minimum score adjustment needed for Vibro or diveXplore to overtake VISIONE under the Figure 5 aggregation. If the observed judge disagreement is far smaller than that margin, the concern does not land; if it is comparable or larger, VISIONE's win is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The report's score ranking depends on a subjective judging step that is never validated. Section 2 states that AVS ground truth cannot be pre-annotated and that QAS answers vary in language and syntax, so live judges manually assessed submissions; 11 judges served. Section 4.1 uses those scores to form the combined ranking in Figure 5, and Section 4.2 omits AVS from the submission plot because it has a comparatively high number of submissions. The paper reports no inter-rater agreement, no adjudication protocol, and no judge-assignment matrix. If judges differ on what counts as correct or on partial credit, the affected scores—especially for AVS—can shift materially, and the paper gives no numeric margins to show that VISIONE's lead exceeds that plausible shift. The archived DRES data makes this testable, but the paper as written does not rule it out.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports the results of the 13th Video Browser Showdown (VBS 2024), held at MMM 2024 in Amsterdam. It describes the competition setup, including the datasets (V3C shards, MVK, and a new laparoscopy collection), the four task types (KISV, KIST, AVS, and QAS), and the use of the Distributed Retrieval Evaluation Server (DRES) with live manual judging for tasks lacking pre-annotated ground truth. The paper lists the 12 participating teams ordered by final rank, with VISIONE first, and presents score distributions for expert and novice sessions, submission statistics, first-submission times, and ranking-over-time plots. The headline claim is that VISIONE achieved the best combined score when the highest-scoring expert and novice participant per team were taken together.","tokens_in":7607,"tokens_out":6447,"duration_ms":60870,"significance":"If accepted as an archival record, the paper is useful to the interactive video retrieval community because it provides a compact snapshot of system performance under a shared evaluation infrastructure, introduces the QAS task for the first time, and makes raw data available through a public archive link. The reported rankings are measured outcomes of a live competition rather than quantities derived from fitted assumptions, so there is no circularity in the central claim. The paper's main transparency strengths are the explicit reliance on DRES and the public data archive; its main weaknesses are the absence of a scoring formula and the lack of any reliability analysis for the manual judging step, both of which limit the reader's ability to verify the final ranking independently.","major_comments":[{"comment":"The scoring formula is not specified anywhere in the paper, although Figure 9 uses the label \"Normalized total score\" and the ranking in Section 3 depends entirely on the scores in Figures 3-5. The authors should state how individual task scores are computed, how they are aggregated across tasks and sessions, and what normalization is applied, or at minimum cite the precise section of a prior detailed VBS report that defines the formula. Without this, the reported score magnitudes and the final ordering cannot be reproduced or interpreted.","section":"Section 4.1 (Figures 3-5)"},{"comment":"The manual judging of AVS and QAS submissions is described as necessary because ground truth cannot be pre-annotated and because answers vary in language and syntax, and the paper states that 11 researchers served as live judges. However, no inter-rater agreement statistics, adjudication protocol, or judge-assignment matrix are reported, and the paper gives no numeric margins in Figure 5 to show that VISIONE's lead exceeds plausible judge variability. Since these manual scores feed directly into the combined ranking, the authors should add a reliability analysis or at least an explicit quantitative sensitivity statement.","section":"Section 2 and Section 4.1"},{"comment":"The combination rule \"highest score per system was combined from either session\" means that teams with more expert or novice participants have more opportunities to contribute a high score, yet the team ranking in Section 3 is based on this combined score. The paper should report the number of expert and novice participants per team, or otherwise justify the rule, because without that information the team ranking may confound system quality with team size.","section":"Section 4.1 (Figure 5)"}],"minor_comments":[{"comment":"The captions of Figures 8 and 10 contain the typo \"Raking\" instead of \"Ranking\".","section":"Figure 8 and Figure 10 captions"},{"comment":"The legend label \"K ISV\" in Figure 4 should be \"KISV\" for consistency with Figure 3 and the rest of the paper.","section":"Figure 4"},{"comment":"The sentence explaining that ad-hoc search tasks are omitted due to their comparatively high number of submissions is clear, but adding one sentence with the order of magnitude of AVS submissions would help the reader interpret the figure.","section":"Section 4.2 (Figure 6)"},{"comment":"The description \"time taken until any participant submitted a first submission for any type of task\" is ambiguous; please clarify whether the plotted distribution is over tasks, participants, or both.","section":"Section 4.2 (Figure 7)"}],"recommendation":"major_revision","confidential_remarks":"The authors are deeply involved in the VBS organization and several of them are affiliated with competing systems, so the paper should include a competing-interests statement. The manuscript is appropriate as a short results report for the multimedia retrieval community, but the missing scoring formula and the unvalidated judging process are load-bearing for the final ranking; I would like to see those addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"VBS 2024 report: it's the annual results paper for the interactive video retrieval competition, so judge it by what it is. New this year: first-time QAS task, a medical laparoscopy dataset, and per-individual scoring instead of per-team aggregate. The paper is transparent about the setup, references each participating system, and points to archived raw data. That is real and useful for the community.\n\nWhat it does well: the task-type breakdown and expert/novice split are clear, the figures show score distributions and rank-over-time, and the archive claim makes the numbers checkable. This is a reproducible benchmark record, not a novel scientific insight, and the paper doesn't claim otherwise.\n\nSoft spots: the scoring formula is never specified, so you can't reconstruct exact scores from the text; you'd need the archive. Manual judging of AVS and QAS is the bigger caveat. Eleven judges assessed submissions with no inter-rater agreement or adjudication protocol reported, and the stress-test note is correct that VISIONE's win rests on unmeasured judge variance. That's a transparency gap, but it's typical for a live competition and the raw data makes it testable. It is not a fatal flaw, but the authors should add these details or at least state why they're omitted.\n\nWho this is for: VBS participants, interactive retrieval researchers, and anyone using benchmark results to compare systems. A serious referee should verify the archive and consistency of figures, but this paper deserves peer review as the official record of a community event. Recommend accept with minor comments on reporting the scoring and judging details.","headline":"A solid, transparent annual VBS results report; the manual judging step is the only real soft spot, but the archived data makes it testable.","tokens_in":8107,"tokens_out":2342,"would_cite":true,"duration_ms":21571,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"VISIONE was the overall best-performing system at the 2024 Video Browser Showdown under the competition's evaluation rules.","keywords":["Video Browser Showdown","interactive video retrieval","known-item search","ad-hoc video search","question answering","evaluation","multimedia retrieval","VBS 2024"],"falsifier":"Recompute the final ranking from the public raw submission logs with an independent ground truth for the manually judged QAS and AVS answers; if VISIONE no longer has the highest combined score under the paper's own aggregation rule, the central claim is false.","tokens_in":7297,"feed_emoji":"🏆","tokens_out":5806,"duration_ms":45323,"temperature":0.7,"pith_summary":"This report presents the results of the 13th Video Browser Showdown, an annual interactive video retrieval competition held at the 2024 International Conference on Multimedia Modeling. It establishes that the VISIONE system achieved the highest overall score among twelve participating teams, based on a combination of each team's best expert and best novice performance. The competition evaluated four task types – known-item search with textual or visual hints, ad-hoc video search, and a new question-answering task – across more than 2,400 hours of video from three datasets, including a novel medical laparoscopy collection. The report's value is as a benchmark: it documents which interactive retrieval approaches worked best under live, time-limited conditions.","feed_headline":"VISIONE wins 2024 Video Browser Showdown","feed_subtitle":"Competition report ranks interactive video retrieval systems across expert and novice sessions on 2,400+ hours of video.","key_machinery":"The machinery that carries the result is the Distributed Retrieval Evaluation Server (DRES), the central coordination and scoring infrastructure that presents tasks, records submissions, and computes scores, together with the explicit aggregation rule that defines the final ranking: for each team, take the best-performing expert participant and the best-performing novice participant, and combine those scores. Manual judging by eleven researchers was used for submissions that could not be automatically checked, specifically for multi-answer ad-hoc search and for question-answering tasks with answers in varying languages. This combined evaluation infrastructure and scoring rule is what makes the reported order of teams a well-defined outcome.","core_discovery":"The paper's central claim is that, under the VBS 2024 evaluation protocol, VISIONE was the overall best-performing video retrieval system. Scores were generated by the Distributed Retrieval Evaluation Server (DRES) and, for answers that could not be automatically verified, by live judges; the final ranking combined the highest score each team achieved in the expert session with the highest score in the novice session. VISIONE placed first in this combined ranking, followed by the other teams in the order listed in the report. The report also shows that novice performance differed substantially from expert performance across systems, and that the newly introduced question-answering task was solved at widely varying rates.","pith_inferences":["One consequence the report leaves implicit is that the best-of aggregation punishes consistency: a system with one strong expert and one strong novice can outrank a system with uniformly strong but lower-peak performance, so the final order is not a measurement of average system quality.","The paper does not discuss statistical significance; the observed score gaps in the figures would need a repeated-competition or per-task variance analysis to show they are not noise.","The medical dataset's novelty suggests a testable extension: rerun a comparable evaluation restricted to laparoscopy videos to see whether the top systems' ranking changes in a homogeneous, domain-specific corpus.","The raw data exports are public, so an independent re-analysis could compare alternative aggregation rules (e.g., per-task averages or median scores) to see how sensitive the final ranking is to the choice of combination."],"forward_implications":["The VBS 2024 ranking gives a concrete baseline: future interactive video retrieval systems can be compared against VISIONE's combined score.","The new question-answering task adds a task type to the evaluation space, so future editions can track how well systems support video-based Q&A.","The separation of expert and novice scores highlights which systems transfer to non-expert users, since novice sessions omitted the textual known-item search task.","The inclusion of a medical laparoscopy dataset extends the evaluation to a specialized domain, broadening the evidence base beyond the V3C and MVK collections.","Because individual participants were scored separately, the results also reveal within-team variability, not just between-team differences."],"supporting_citations":[{"why":"Supplies the Distributed Retrieval Evaluation Server that generated all the scores and coordinated the tasks.","marker":"[4]"},{"why":"Identifies the VISIONE system that the report ranks first, tying the winning score to a specific retrieval tool.","marker":"[8]"},{"why":"Provides the V3C video collection, one of the two main evaluation datasets with 2,300 hours of content.","marker":"[5]"},{"why":"Provides the MVK marine video dataset used for a second, smaller evaluation collection.","marker":"[6]"},{"why":"Establishes the evaluation methodology and setup carried over from the previous year's VBS competition.","marker":"[1]"},{"why":"Defines the task category space that motivates the choice of task types used in VBS 2024.","marker":"[7]"}],"fun_headline_variants":["VISIONE tops combined expert-novice ranking at VBS 2024","VISIONE wins 2024 VBS with best expert and novice scores","VISIONE leads VBS 2024 combined ranking of video search","VBS 2024: VISIONE first in combined expert-novice results","2024 Video Browser Showdown: VISIONE takes overall win"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the evaluation server and the live judges scored every submission correctly and consistently, and that combining each team's best expert and novice score into a single ranking is a fair way to decide the winner.","fun_headline_variants_meta":{"raw":{"variants":["VISIONE tops combined expert-novice ranking at VBS 2024","VISIONE wins 2024 VBS with best expert and novice scores","VISIONE leads VBS 2024 combined ranking of video search","VBS 2024: VISIONE first in combined expert-novice results","2024 Video Browser Showdown: VISIONE takes overall win"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000644,"raw_usage":{"total_tokens":2833,"prompt_tokens":687,"completion_tokens":2146,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":303,"completion_tokens_details":{"reasoning_tokens":2045}},"tokens_in":303,"tokens_out":2146,"duration_ms":13275,"temperature":1.0,"reasoning_tokens":2045,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:51:15.545840+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the final ranking from the public raw submission logs with an independent ground truth for the manually judged QAS and AVS answers; if VISIONE no longer has the highest combined score under the paper's own aggregation rule, the central claim is false.","supporting_citations":[{"cited_title":"Sauter, R","cited_arxiv_id":null,"evidence_quote":"Supplies the Distributed Retrieval Evaluation Server that generated all the scores and coordinated the tasks."},{"cited_title":"Rossetto, H","cited_arxiv_id":null,"evidence_quote":"Provides the V3C video collection, one of the two main evaluation datasets with 2,300 hours of content."},{"cited_title":"Truong, T.-A","cited_arxiv_id":null,"evidence_quote":"Provides the MVK marine video dataset used for a second, smaller evaluation collection."},{"cited_title":"Lokoc, W","cited_arxiv_id":null,"evidence_quote":"Defines the task category space that motivates the choice of task types used in VBS 2024."}],"review_version":1}