Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Results of the 2024 Video Browser Showdown

T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read VISIONE was the overall best-performing system at the 2024 Video Browser Showdown under the competition's evaluation rules.

desk verdict A solid, transparent annual VBS results report; the manual judging step is the only real soft spot, but the archived data makes it testable. read the letter →

arxiv 2502.15683 v1 pith:7LJEDRJ4 submitted 2024-12-13 cs.MM cs.IR

classification cs.MMcs.IR
keywords VideoBrowserShowdowninteractiveretrievalknown-itemsearchad-hocquestionansweringevaluationmultimediaVBS2024
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This report presents the results of the 13th Video Browser Showdown, an annual interactive video retrieval competition held at the 2024 International Conference on Multimedia Modeling. It establishes that the VISIONE system achieved the highest overall score among twelve participating teams, based on a combination of each team's best expert and best novice performance. The competition evaluated four task types – known-item search with textual or visual hints, ad-hoc video search, and a new question-answering task – across more than 2,400 hours of video from three datasets, including a novel medical laparoscopy collection. The report's value is as a benchmark: it documents which interactive retrieval approaches worked best under live, time-limited conditions.

What carries the argument

The machinery that carries the result is the Distributed Retrieval Evaluation Server (DRES), the central coordination and scoring infrastructure that presents tasks, records submissions, and computes scores, together with the explicit aggregation rule that defines the final ranking: for each team, take the best-performing expert participant and the best-performing novice participant, and combine those scores. Manual judging by eleven researchers was used for submissions that could not be automatically checked, specifically for multi-answer ad-hoc search and for question-answering tasks with answers in varying languages. This combined evaluation infrastructure and scoring rule is what makes the reported order of teams a well-defined outcome.

What would settle it

Recompute the final ranking from the public raw submission logs with an independent ground truth for the manually judged QAS and AVS answers; if VISIONE no longer has the highest combined score under the paper's own aggregation rule, the central claim is false.

Watch

Extended reading notes

Core claim

The paper's central claim is that, under the VBS 2024 evaluation protocol, VISIONE was the overall best-performing video retrieval system. Scores were generated by the Distributed Retrieval Evaluation Server (DRES) and, for answers that could not be automatically verified, by live judges; the final ranking combined the highest score each team achieved in the expert session with the highest score in the novice session. VISIONE placed first in this combined ranking, followed by the other teams in the order listed in the report. The report also shows that novice performance differed substantially from expert performance across systems, and that the newly introduced question-answering task was solved at widely varying rates.

Load-bearing premise

The load-bearing premise is that the evaluation server and the live judges scored every submission correctly and consistently, and that combining each team's best expert and novice score into a single ranking is a fair way to decide the winner.

Editorial extensions

If this is right

  • The VBS 2024 ranking gives a concrete baseline: future interactive video retrieval systems can be compared against VISIONE's combined score.
  • The new question-answering task adds a task type to the evaluation space, so future editions can track how well systems support video-based Q&A.
  • The separation of expert and novice scores highlights which systems transfer to non-expert users, since novice sessions omitted the textual known-item search task.
  • The inclusion of a medical laparoscopy dataset extends the evaluation to a specialized domain, broadening the evidence base beyond the V3C and MVK collections.
  • Because individual participants were scored separately, the results also reveal within-team variability, not just between-team differences.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One consequence the report leaves implicit is that the best-of aggregation punishes consistency: a system with one strong expert and one strong novice can outrank a system with uniformly strong but lower-peak performance, so the final order is not a measurement of average system quality.
  • The paper does not discuss statistical significance; the observed score gaps in the figures would need a repeated-competition or per-task variance analysis to show they are not noise.
  • The medical dataset's novelty suggests a testable extension: rerun a comparable evaluation restricted to laparoscopy videos to see whether the top systems' ranking changes in a homogeneous, domain-specific corpus.
  • The raw data exports are public, so an independent re-analysis could compare alternative aggregation rules (e.g., per-task averages or median scores) to see how sensitive the final ranking is to the choice of combination.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This manuscript reports the results of the 13th Video Browser Showdown (VBS 2024), held at MMM 2024 in Amsterdam. It describes the competition setup, including the datasets (V3C shards, MVK, and a new laparoscopy collection), the four task types (KISV, KIST, AVS, and QAS), and the use of the Distributed Retrieval Evaluation Server (DRES) with live manual judging for tasks lacking pre-annotated ground truth. The paper lists the 12 participating teams ordered by final rank, with VISIONE first, and presents score distributions for expert and novice sessions, submission statistics, first-submission times, and ranking-over-time plots. The headline claim is that VISIONE achieved the best combined score when the highest-scoring expert and novice participant per team were taken together.

Significance. If accepted as an archival record, the paper is useful to the interactive video retrieval community because it provides a compact snapshot of system performance under a shared evaluation infrastructure, introduces the QAS task for the first time, and makes raw data available through a public archive link. The reported rankings are measured outcomes of a live competition rather than quantities derived from fitted assumptions, so there is no circularity in the central claim. The paper's main transparency strengths are the explicit reliance on DRES and the public data archive; its main weaknesses are the absence of a scoring formula and the lack of any reliability analysis for the manual judging step, both of which limit the reader's ability to verify the final ranking independently.

major comments (3)
  1. [Section 4.1 (Figures 3-5)] The scoring formula is not specified anywhere in the paper, although Figure 9 uses the label "Normalized total score" and the ranking in Section 3 depends entirely on the scores in Figures 3-5. The authors should state how individual task scores are computed, how they are aggregated across tasks and sessions, and what normalization is applied, or at minimum cite the precise section of a prior detailed VBS report that defines the formula. Without this, the reported score magnitudes and the final ordering cannot be reproduced or interpreted.
  2. [Section 2 and Section 4.1] The manual judging of AVS and QAS submissions is described as necessary because ground truth cannot be pre-annotated and because answers vary in language and syntax, and the paper states that 11 researchers served as live judges. However, no inter-rater agreement statistics, adjudication protocol, or judge-assignment matrix are reported, and the paper gives no numeric margins in Figure 5 to show that VISIONE's lead exceeds plausible judge variability. Since these manual scores feed directly into the combined ranking, the authors should add a reliability analysis or at least an explicit quantitative sensitivity statement.
  3. [Section 4.1 (Figure 5)] The combination rule "highest score per system was combined from either session" means that teams with more expert or novice participants have more opportunities to contribute a high score, yet the team ranking in Section 3 is based on this combined score. The paper should report the number of expert and novice participants per team, or otherwise justify the rule, because without that information the team ranking may confound system quality with team size.
minor comments (4)
  1. [Figure 8 and Figure 10 captions] The captions of Figures 8 and 10 contain the typo "Raking" instead of "Ranking".
  2. [Figure 4] The legend label "K ISV" in Figure 4 should be "KISV" for consistency with Figure 3 and the rest of the paper.
  3. [Section 4.2 (Figure 6)] The sentence explaining that ad-hoc search tasks are omitted due to their comparatively high number of submissions is clear, but adding one sentence with the order of magnitude of AVS submissions would help the reader interpret the figure.
  4. [Section 4.2 (Figure 7)] The description "time taken until any participant submitted a first submission for any type of task" is ambiguous; please clarify whether the plotted distribution is over tasks, participants, or both.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper reports measured competition scores; rankings are aggregations of observed data, not derived predictions.

full rationale

This paper is a results report for a live competition. The scores shown in Figures 3-5 were generated by the Distributed Retrieval Evaluation Server (DRES) and live judges, as stated in Section 2, and the ranking of teams in Section 3 is simply the ordering of those measured scores. There is no fitted parameter, no model-derived quantity, and no theoretical prediction that could reduce to its own inputs. The combination rule in Section 4.1 ('the highest score per system was combined from either session') is an explicit aggregation procedure applied to observed scores, not a hidden assumption that manufactures the ranking. The authors' prior involvement in VBS infrastructure and their self-citations to previous VBS reports and the DRES system are contextual and do not carry the load-bearing argument; the load-bearing evidence is the archived evaluation data. The skeptical concern about inter-judge consistency for AVS and QAS submissions is a validity or measurement-quality issue, not a circularity issue: even if judge variance changed the scores, the reported numbers would still be observed measurements rather than quantities derived from the claims they support. Accordingly, no circular step can be identified, and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The report relies on the integrity of the evaluation server and judges; no free parameters or invented entities are introduced.

assumptions (2)
  • domain assumption The Distributed Retrieval Evaluation Server (DRES) records and scores submissions correctly.
    The paper states in Section 2 that data was generated by DRES and in Section 4.1 that scores are shown; the correctness of the reported results depends on this infrastructure.
  • domain assumption Live judges correctly assess submissions for tasks without complete ground truth.
    Section 2 notes that AVS and QAS answers required manual judging; incorrect judgments would change scores and rankings.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Results of the 2024 Video Browser Showdown." pith.science (2026). https://pith.science/paper/7LJEDRJ4

@misc{pith2026250215683,
  author       = {Pith},
  title        = {Pith review of: Results of the 2024 Video Browser Showdown},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7LJEDRJ4}},
  note         = {Machine review of arXiv:2502.15683}
}
read the original abstract

This report presents the results of the 13th Video Browser Showdown, held at the 2024 International Conference on Multimedia Modeling on the 29th of January 2024 in Amsterdam, the Netherlands.

Figures

Figures reproduced from arXiv: 2502.15683 by the authors.

Figure 1
Figure 1. Setup of VBS 2024: a u-shape arrangement of teams around the projector showing the evaluation server display (tasks and results). 3. Teams In 2024, 12 teams participated in VBS. In contrast to previous years, team members were scored individually rather than in aggregate. The following lists the participating teams, ordered by their final rank in the evaluation. 1. VISIONE [8] 2. Vibro [9] 3. PraK [10] 4. ViewsInsig… view at source ↗
Figure 2
Figure 2. Live judging of submissions by a team of on-site and offline judges for tasks without (complete) ground-truth. 4. Results The following illustrates the scores, including their development over time, as well as the general properties of the submissions made by all participants during the evaluation. The evaluation was split into two parts: first, the developers of the participating systems operated their systems them… view at source ↗
Figure 3
Figure 3. Scores of the expert tasks grouped by participant Team VISIONE1 Vibro2 diveXplore1 PraK4 Exquisitor1 ViewsInsight3 VIREO1 VERGE1 vitrivr4 TalkSee2 PraK2 Exquisitor2 vitrivr2 ViewsInsight2 VIREO2 TalkSee1 vitrivr3 PraK3 VISIONE2 VERGE2 vitrivr1 vitrivr-VR1 VISIONE3 ViewsInsight1 VideoCLIP2 VideoCLIP1 PraK1 diveXplore2 Vibro1 VISIONE4 vitrivr-VR2 PraK5 S core 0 500 1000 1500 2000 2500 Novice Tasks Task Type AVS KISV Q… view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Scores of the novice tasks grouped by participant Team VISIONE Vibro diveXplore PraK ViewsInsight TalkSee VIREO vitrivr vitrivr-VR Exquisitor VERGE VideoCLIP Scor e 0 1000 2000 3000 4000 5000 6000 Combined Scores Task Type Expert Novice [PITH_FULL_IMAGE:figures/full_f…
Figure 5
Figure 5. Figure 5: Combined scores, using the best-performing expert and novice per team [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: shows the number of correct and incorrect submissions per type of task and participant. Ad-hoc search tasks are omitted due to their comparatively high number of submissions. Team VISIONE1 Vibro1 VISIONE4 VISIONE2 PraK5 PraK2 Vibro2 PraK4 ViewsInsight1 VISIONE3 ViewsIn…
Figure 7
Figure 7. Figure 7: shows the distribution of the time taken until any participant submitted a first submission for any type of task. This serves as a rough estimate of the query processing efficiency of the different systems. Team VISIONE1 Vibro1 VISIONE4 VISIONE2 PraK5 PraK2 Vibro2 PraK…
Figure 8
Figure 8. Figure 8: shows the ranking of the expert participants after every task in the evaluation. Task 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 T e a m Rank 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 Expert rank over time T…
Figure 9
Figure 9. Figure 9: shows the score of the expert participants after each task in the evaluation. Task 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 Normalize d total sc ore 0 500 1000 1500 2000 2500 3000 3500 Expert score over time Team Vibro1 VISIONE1 VISIONE4 PraK5 PraK2 …
Figure 10
Figure 10. Figure 10: shows the ranking of novice participants after each task in the evaluation. Task 1 2 3 4 5 6 7 8 T e a m Rank 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 Novice rank over time Team VISIONE1 diveXplore1 Vibro2 PraK4 Exquisitor…
Figure 11
Figure 11. Figure 11: shows the score of novice participants after each task in the evaluation. Task 1 2 3 4 5 6 7 8 Normalize d total sc ore 0 500 1000 1500 2000 2500 Novice score over time Team VISIONE1 diveXplore1 Vibro2 PraK4 Exquisitor1 ViewsInsight3 VIREO1 VERGE1 vitrivr4 TalkSee2 […

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning

    cs.CV 2025-08 reject novelty 6.0 of 10

    A new marine wildlife video dataset with clip-level captions and grounded segmentation masks, benchmarked on captioning, grounding, and video generation.

Reference graph

Works this paper leans on

17 extracted references · 12 canonical work pages · cited by 1 Pith paper

  1. [1]

    Vadicamo, R

    L. Vadicamo, R. Arnold, W. Bailer, F. Carrara, C. Gurrin, N. Hezel, X. Li, J. Lokoc, S. Lubos, Z. Ma, N. Messina, T. Nguyen, L. Peska, L. Rossetto, L. Sauter, K. Schöffmann, F. Spiess, M. Tran, S. Vrochidis, Evaluating performance and trends in interactive video retrieval: Insights from the 12th VBS competition, IEEE Access 12 (2024) 79342–79366. doi: 10....

  2. [2]

    Lokoc, S

    J. Lokoc, S. Andreadis, W. Bailer, A. Duane, C. Gurrin, Z. Ma, N. Messina, T. Nguyen, L. Peska, L. Rossetto, L. Sauter, K. Schall, K. Schoeffmann, O. S. Khan, F. Spiess, L. Vadicamo, S. Vrochidis, Interactive video retrieval in the age of effective joint embedding deep models: lessons from the 11th VBS, Multim. Syst. 29 (2023) 3481–3504. doi: 10.1007/ S00...

  3. [3]

    Heller, V

    S. Heller, V. Gsteiger, W. Bailer, C. Gurrin, B. Þ. Jónsson, J. Lokoc, A. Leibetseder, F. Mejzlík, L. Peska, L. Rossetto, K. Schall, K. Schoeffmann, H. Schuldt, F. Spiess, L. Tran, L. Vadicamo, P. Veselý, S. Vrochidis, J. Wu, Interactive video retrieval evaluation at a distance: comparing sixteen interactive video search systems in a remote setting at the...

  4. [4]

    Sauter, R

    L. Sauter, R. Gasser, H. Schuldt, A. Bernstein, L. Rossetto, Performance evaluation in multimedia retrieval, ACM Trans. Multimedia Comput. Commun. Appl. (2024). doi: 10. 1145/3678881

  5. [5]

    Rossetto, H

    L. Rossetto, H. Schuldt, G. Awad, A. A. Butt, V3C - A research video collection, in: International Conference on Multimedia Modeling, Springer, 2019, pp. 349–360. doi:10. 1007/978-3-030-05710-7\_29

  6. [6]

    Truong, T.-A

    Q.-T. Truong, T.-A. Vu, T.-S. Ha, J. Lokoč, Y. H. W. Tim, A. Joneja, S.-K. Yeung, Marine video kit: A new marine video dataset for content-based analysis and retrieval, in: MultiMedia Modeling - 29th International Conference, MMM 2023, Bergen, Norway, January 9-12, 2023, Springer, 2023. doi:10.1007/978-3-031-27077-2_42

  7. [7]

    Lokoc, W

    J. Lokoc, W. Bailer, K. U. Barthel, C. Gurrin, S. Heller, B. Þ. Jónsson, L. Peska, L. Rossetto, K. Schoeffmann, L. Vadicamo, S. Vrochidis, J. Wu, A task category space for user-centric comparative multimedia search evaluations, in: MultiMedia Modeling - 28th International Conference, MMM 2022, Phu Quoc, Vietnam, June 6-10, 2022, Proceedings, Part I, volum...

  8. [8]

    Amato, P

    G. Amato, P. Bolettieri, F. Carrara, F. Falchi, C. Gennaro, N. Messina, L. Vadicamo, C. Vairo, VISIONE 5.0: Enhanced user interface and AI models for VBS2024, in: S. Rudinac, A. Han- jalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Yamakata (Eds.), MultiMedia Modeling - 30th International Conference, MMM 2024, Amsterdam, The Netherlands, Jan- u...

Show all 17 references
  1. [11]

    Vuong, V

    G. Vuong, V. Ho, T. Nguyen-Dang, X. Thai, T. Le, M. Pham, V. Ninh, C. Gurrin, M. Tran, Viewsinsight: Enhancing video retrieval for VBS 2024 with a user-friendly interaction mechanism, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Yamakata (E...

  2. [12]

    Schoeffmann, S

    K. Schoeffmann, S. Nasirihaghighi, Divexplore at the video browser showdown 2024, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Ya- makata (Eds.), MultiMedia Modeling - 30th International Conference, MMM 2024, Am- sterdam, The Netherlands, J...

  3. [13]

    G. Gu, Z. Wu, J. He, L. Song, Z. Wang, C. Liang, Talksee: Interactive video retrieval engine using large language model, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Yamakata (Eds.), MultiMedia Modeling - 30th International Conference, MMM ...

  4. [14]

    Z. Ma, J. Wu, C. W. Ngo, Leveraging llms and generative models for interactive known-item video search, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Yamakata (Eds.), MultiMedia Modeling - 30th International Conference, MMM 2024, Amsterdam, ...

  5. [15]

    Spiess, L

    F. Spiess, L. Rossetto, H. Schuldt, Exploring multimedia vector spaces with vitrivr-vr, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Ya- makata (Eds.), MultiMedia Modeling - 30th International Conference, MMM 2024, Am- sterdam, The Netherla...

  6. [16]

    Gasser, R

    R. Gasser, R. Arnold, F. Faber, H. Schuldt, R. Waltenspül, L. Rossetto, A new retrieval engine for vitrivr, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Yamakata (Eds.), MultiMedia Modeling - 30th International Conference, MMM 2024, Amsterd...

  7. [17]

    Pantelidis, M

    N. Pantelidis, M. Pegia, D. Galanopoulos, K. Apostolidis, K. Stavrothanasopoulos, A. Moumtzidou, K. Gkountakos, I. Gialampoukidis, S. Vrochidis, V. Mezaris, I. Kompatsiaris, B. Þ. Jónsson, VERGE in VBS 2024, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Wor- ring, B. Þ. Jónss...

  8. [18]

    Nguyen, L

    T. Nguyen, L. M. Quang, G. Healy, B. T. Nguyen, C. Gurrin, Videoclip 2.0: An interactive clip-based video retrieval system for novice users at VBS2024, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Yamakata (Eds.), MultiMedia Modeling - 30th...

  9. [19]

    O. S. Khan, H. Zhu, U. Sharma, E. Kanoulas, S. Rudinac, B. Þ. Jónsson, Exquisitor at the video browser showdown 2024: Relevance feedback meets conversational search, in: S. Rudinac, A. Hanjalic, C. C. S. Liem, M. Worring, B. Þ. Jónsson, B. Liu, Y. Yamakata (Eds.), MultiMedia M...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.