REVIEW 3 major objections 4 minor 1 cited by
Results of the 2024 Video Browser Showdown
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read VISIONE was the overall best-performing system at the 2024 Video Browser Showdown under the competition's evaluation rules.
desk verdict A solid, transparent annual VBS results report; the manual judging step is the only real soft spot, but the archived data makes it testable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery that carries the result is the Distributed Retrieval Evaluation Server (DRES), the central coordination and scoring infrastructure that presents tasks, records submissions, and computes scores, together with the explicit aggregation rule that defines the final ranking: for each team, take the best-performing expert participant and the best-performing novice participant, and combine those scores. Manual judging by eleven researchers was used for submissions that could not be automatically checked, specifically for multi-answer ad-hoc search and for question-answering tasks with answers in varying languages. This combined evaluation infrastructure and scoring rule is what makes the reported order of teams a well-defined outcome.
What would settle it
Recompute the final ranking from the public raw submission logs with an independent ground truth for the manually judged QAS and AVS answers; if VISIONE no longer has the highest combined score under the paper's own aggregation rule, the central claim is false.
Extended reading notes
Core claim
The paper's central claim is that, under the VBS 2024 evaluation protocol, VISIONE was the overall best-performing video retrieval system. Scores were generated by the Distributed Retrieval Evaluation Server (DRES) and, for answers that could not be automatically verified, by live judges; the final ranking combined the highest score each team achieved in the expert session with the highest score in the novice session. VISIONE placed first in this combined ranking, followed by the other teams in the order listed in the report. The report also shows that novice performance differed substantially from expert performance across systems, and that the newly introduced question-answering task was solved at widely varying rates.
Load-bearing premise
The load-bearing premise is that the evaluation server and the live judges scored every submission correctly and consistently, and that combining each team's best expert and novice score into a single ranking is a fair way to decide the winner.
Editorial extensions
If this is right
- The VBS 2024 ranking gives a concrete baseline: future interactive video retrieval systems can be compared against VISIONE's combined score.
- The new question-answering task adds a task type to the evaluation space, so future editions can track how well systems support video-based Q&A.
- The separation of expert and novice scores highlights which systems transfer to non-expert users, since novice sessions omitted the textual known-item search task.
- The inclusion of a medical laparoscopy dataset extends the evaluation to a specialized domain, broadening the evidence base beyond the V3C and MVK collections.
- Because individual participants were scored separately, the results also reveal within-team variability, not just between-team differences.
Reading between the lines
- One consequence the report leaves implicit is that the best-of aggregation punishes consistency: a system with one strong expert and one strong novice can outrank a system with uniformly strong but lower-peak performance, so the final order is not a measurement of average system quality.
- The paper does not discuss statistical significance; the observed score gaps in the figures would need a repeated-competition or per-task variance analysis to show they are not noise.
- The medical dataset's novelty suggests a testable extension: rerun a comparable evaluation restricted to laparoscopy videos to see whether the top systems' ranking changes in a homogeneous, domain-specific corpus.
- The raw data exports are public, so an independent re-analysis could compare alternative aggregation rules (e.g., per-task averages or median scores) to see how sensitive the final ranking is to the choice of combination.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript reports the results of the 13th Video Browser Showdown (VBS 2024), held at MMM 2024 in Amsterdam. It describes the competition setup, including the datasets (V3C shards, MVK, and a new laparoscopy collection), the four task types (KISV, KIST, AVS, and QAS), and the use of the Distributed Retrieval Evaluation Server (DRES) with live manual judging for tasks lacking pre-annotated ground truth. The paper lists the 12 participating teams ordered by final rank, with VISIONE first, and presents score distributions for expert and novice sessions, submission statistics, first-submission times, and ranking-over-time plots. The headline claim is that VISIONE achieved the best combined score when the highest-scoring expert and novice participant per team were taken together.
Significance. If accepted as an archival record, the paper is useful to the interactive video retrieval community because it provides a compact snapshot of system performance under a shared evaluation infrastructure, introduces the QAS task for the first time, and makes raw data available through a public archive link. The reported rankings are measured outcomes of a live competition rather than quantities derived from fitted assumptions, so there is no circularity in the central claim. The paper's main transparency strengths are the explicit reliance on DRES and the public data archive; its main weaknesses are the absence of a scoring formula and the lack of any reliability analysis for the manual judging step, both of which limit the reader's ability to verify the final ranking independently.
major comments (3)
- [Section 4.1 (Figures 3-5)] The scoring formula is not specified anywhere in the paper, although Figure 9 uses the label "Normalized total score" and the ranking in Section 3 depends entirely on the scores in Figures 3-5. The authors should state how individual task scores are computed, how they are aggregated across tasks and sessions, and what normalization is applied, or at minimum cite the precise section of a prior detailed VBS report that defines the formula. Without this, the reported score magnitudes and the final ordering cannot be reproduced or interpreted.
- [Section 2 and Section 4.1] The manual judging of AVS and QAS submissions is described as necessary because ground truth cannot be pre-annotated and because answers vary in language and syntax, and the paper states that 11 researchers served as live judges. However, no inter-rater agreement statistics, adjudication protocol, or judge-assignment matrix are reported, and the paper gives no numeric margins in Figure 5 to show that VISIONE's lead exceeds plausible judge variability. Since these manual scores feed directly into the combined ranking, the authors should add a reliability analysis or at least an explicit quantitative sensitivity statement.
- [Section 4.1 (Figure 5)] The combination rule "highest score per system was combined from either session" means that teams with more expert or novice participants have more opportunities to contribute a high score, yet the team ranking in Section 3 is based on this combined score. The paper should report the number of expert and novice participants per team, or otherwise justify the rule, because without that information the team ranking may confound system quality with team size.
minor comments (4)
- [Figure 8 and Figure 10 captions] The captions of Figures 8 and 10 contain the typo "Raking" instead of "Ranking".
- [Figure 4] The legend label "K ISV" in Figure 4 should be "KISV" for consistency with Figure 3 and the rest of the paper.
- [Section 4.2 (Figure 6)] The sentence explaining that ad-hoc search tasks are omitted due to their comparatively high number of submissions is clear, but adding one sentence with the order of magnitude of AVS submissions would help the reader interpret the figure.
- [Section 4.2 (Figure 7)] The description "time taken until any participant submitted a first submission for any type of task" is ambiguous; please clarify whether the plotted distribution is over tasks, participants, or both.
Circularity Check
No circularity: the paper reports measured competition scores; rankings are aggregations of observed data, not derived predictions.
full rationale
This paper is a results report for a live competition. The scores shown in Figures 3-5 were generated by the Distributed Retrieval Evaluation Server (DRES) and live judges, as stated in Section 2, and the ranking of teams in Section 3 is simply the ordering of those measured scores. There is no fitted parameter, no model-derived quantity, and no theoretical prediction that could reduce to its own inputs. The combination rule in Section 4.1 ('the highest score per system was combined from either session') is an explicit aggregation procedure applied to observed scores, not a hidden assumption that manufactures the ranking. The authors' prior involvement in VBS infrastructure and their self-citations to previous VBS reports and the DRES system are contextual and do not carry the load-bearing argument; the load-bearing evidence is the archived evaluation data. The skeptical concern about inter-judge consistency for AVS and QAS submissions is a validity or measurement-quality issue, not a circularity issue: even if judge variance changed the scores, the reported numbers would still be observed measurements rather than quantities derived from the claims they support. Accordingly, no circular step can be identified, and the appropriate score is 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The Distributed Retrieval Evaluation Server (DRES) records and scores submissions correctly.
- domain assumption Live judges correctly assess submissions for tasks without complete ground truth.
Cite this review
Pith. "Pith review of Results of the 2024 Video Browser Showdown." pith.science (2026). https://pith.science/paper/7LJEDRJ4
@misc{pith2026250215683,
author = {Pith},
title = {Pith review of: Results of the 2024 Video Browser Showdown},
year = {2026},
howpublished = {\url{https://pith.science/paper/7LJEDRJ4}},
note = {Machine review of arXiv:2502.15683}
}
read the original abstract
This report presents the results of the 13th Video Browser Showdown, held at the 2024 International Conference on Multimedia Modeling on the 29th of January 2024 in Amsterdam, the Netherlands.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 1 Pith paper
-
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
A new marine wildlife video dataset with clip-level captions and grounded segmentation masks, benchmarked on captioning, grounding, and video generation.
Reference graph
Works this paper leans on
-
[1]
L. Vadicamo, R. Arnold, W. Bailer, F. Carrara, C. Gurrin, N. Hezel, X. Li, J. Lokoc, S. Lubos, Z. Ma, N. Messina, T. Nguyen, L. Peska, L. Rossetto, L. Sauter, K. Schöffmann, F. Spiess, M. Tran, S. Vrochidis, Evaluating performance and trends in interactive video retrieval: Insights from the 12th VBS competition, IEEE Access 12 (2024) 79342–79366. doi: 10....
arXiv 2024
-
[2]
J. Lokoc, S. Andreadis, W. Bailer, A. Duane, C. Gurrin, Z. Ma, N. Messina, T. Nguyen, L. Peska, L. Rossetto, L. Sauter, K. Schall, K. Schoeffmann, O. S. Khan, F. Spiess, L. Vadicamo, S. Vrochidis, Interactive video retrieval in the age of effective joint embedding deep models: lessons from the 11th VBS, Multim. Syst. 29 (2023) 3481–3504. doi: 10.1007/ S00...
work page 2023
-
[3]
S. Heller, V. Gsteiger, W. Bailer, C. Gurrin, B. Þ. Jónsson, J. Lokoc, A. Leibetseder, F. Mejzlík, L. Peska, L. Rossetto, K. Schall, K. Schoeffmann, H. Schuldt, F. Spiess, L. Tran, L. Vadicamo, P. Veselý, S. Vrochidis, J. Wu, Interactive video retrieval evaluation at a distance: comparing sixteen interactive video search systems in a remote setting at the...
- [4]
-
[5]
L. Rossetto, H. Schuldt, G. Awad, A. A. Butt, V3C - A research video collection, in: International Conference on Multimedia Modeling, Springer, 2019, pp. 349–360. doi:10. 1007/978-3-030-05710-7\_29
work page 2019
-
[6]
Q.-T. Truong, T.-A. Vu, T.-S. Ha, J. Lokoč, Y. H. W. Tim, A. Joneja, S.-K. Yeung, Marine video kit: A new marine video dataset for content-based analysis and retrieval, in: MultiMedia Modeling - 29th International Conference, MMM 2023, Bergen, Norway, January 9-12, 2023, Springer, 2023. doi:10.1007/978-3-031-27077-2_42
-
[7]
J. Lokoc, W. Bailer, K. U. Barthel, C. Gurrin, S. Heller, B. Þ. Jónsson, L. Peska, L. Rossetto, K. Schoeffmann, L. Vadicamo, S. Vrochidis, J. Wu, A task category space for user-centric comparative multimedia search evaluations, in: MultiMedia Modeling - 28th International Conference, MMM 2022, Phu Quoc, Vietnam, June 6-10, 2022, Proceedings, Part I, volum...
work page 2022
-
[8]
G. Amato, P. Bolettieri, F. Carrara, F. Falchi, C. Gennaro, N. Messina, L. Vadicamo, C. Vairo, VISIONE 5.0: Enhanced user interface and AI models for VBS2024, in: S. Rudinac, A. Han- jalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Yamakata (Eds.), MultiMedia Modeling - 30th International Conference, MMM 2024, Amsterdam, The Netherlands, Jan- u...
doi:10.1007/978 2024
Show all 17 references
-
[11]
Vuong, V
G. Vuong, V. Ho, T. Nguyen-Dang, X. Thai, T. Le, M. Pham, V. Ninh, C. Gurrin, M. Tran, Viewsinsight: Enhancing video retrieval for VBS 2024 with a user-friendly interaction mechanism, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Yamakata (E...
2024
-
[12]
Schoeffmann, S
K. Schoeffmann, S. Nasirihaghighi, Divexplore at the video browser showdown 2024, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Ya- makata (Eds.), MultiMedia Modeling - 30th International Conference, MMM 2024, Am- sterdam, The Netherlands, J...
2024
-
[13]
G. Gu, Z. Wu, J. He, L. Song, Z. Wang, C. Liang, Talksee: Interactive video retrieval engine using large language model, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Yamakata (Eds.), MultiMedia Modeling - 30th International Conference, MMM ...
2024 doi
-
[14]
Z. Ma, J. Wu, C. W. Ngo, Leveraging llms and generative models for interactive known-item video search, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Yamakata (Eds.), MultiMedia Modeling - 30th International Conference, MMM 2024, Amsterdam, ...
2024
-
[15]
Spiess, L
F. Spiess, L. Rossetto, H. Schuldt, Exploring multimedia vector spaces with vitrivr-vr, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Ya- makata (Eds.), MultiMedia Modeling - 30th International Conference, MMM 2024, Am- sterdam, The Netherla...
2024
-
[16]
Gasser, R
R. Gasser, R. Arnold, F. Faber, H. Schuldt, R. Waltenspül, L. Rossetto, A new retrieval engine for vitrivr, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Yamakata (Eds.), MultiMedia Modeling - 30th International Conference, MMM 2024, Amsterd...
2024
-
[17]
Pantelidis, M
N. Pantelidis, M. Pegia, D. Galanopoulos, K. Apostolidis, K. Stavrothanasopoulos, A. Moumtzidou, K. Gkountakos, I. Gialampoukidis, S. Vrochidis, V. Mezaris, I. Kompatsiaris, B. Þ. Jónsson, VERGE in VBS 2024, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Wor- ring, B. Þ. Jónss...
2024
-
[18]
Nguyen, L
T. Nguyen, L. M. Quang, G. Healy, B. T. Nguyen, C. Gurrin, Videoclip 2.0: An interactive clip-based video retrieval system for novice users at VBS2024, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Yamakata (Eds.), MultiMedia Modeling - 30th...
2024 doi
-
[19]
O. S. Khan, H. Zhu, U. Sharma, E. Kanoulas, S. Rudinac, B. Þ. Jónsson, Exquisitor at the video browser showdown 2024: Relevance feedback meets conversational search, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Yamakata (Eds.), MultiMedia M...
2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.