Pith. sign in

REVIEW 3 major objections 6 minor 98 references

A three-pillar framework—datasets, probabilistic evaluation, and deployment scenarios—aims to make video event-detection methods comparable without the biases of small, narrow benchmarks.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 19:39 UTC pith:GNUQE5YX

load-bearing objection Solid three-pillar methodology paper: new multi-environment datasets and scenario checklist are the real additions; the ranking/Tile math is mostly restated prior work and the crisp-classification premise is a scope condition, not a flaw. the 3 major comments →

arxiv 2607.04372 v1 pith:GNUQE5YX submitted 2026-07-05 cs.CV

Event Detection in Videos: A Framework for the Development of New Methods

classification cs.CV
keywords event detectionvideo surveillancebackground subtractionperformance evaluationranking scoresTilesapplication scenarioslarge-scale datasets
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Event detection in video (pixel, frame, or clip level) has long been compared with small, application-narrow datasets and simple average rankings, which the authors say has biased both progress and claims. This paper argues that fair development and comparison require an explicit three-pillar framework. First, videos are organized by environment and modality and tagged by visual challenge so researchers can select balanced or targeted subsets; the authors also release two new datasets (a real public-IP-camera set and a synthetic urban-crossroad set) with public and private splits. Second, every detection task is cast as a two-class crisp classification whose performance is a probability measure over the four outcomes true-negative, false-positive, false-negative, and true-positive; those performances are averaged by mixing data sources, analyzed with a continuum of importance-weighted ranking scores visualized as “Tiles,” and used to produce stable method rankings. Third, an “application scenario” must declare which data, prior knowledge, and evaluation protocol a method may use so that results remain comparable. Together the pillars are meant to remove under-specification, chance-level inflation, and hidden training advantages from the literature.

Core claim

A rigorous framework for new video event-detection methods rests on three pillars: (1) a hierarchically structured, multi-environment, tag-annotated large-scale collection of videos that includes two new public/private datasets; (2) a probabilistic evaluation pipeline that represents performance as a probability measure on the four crisp outcomes, summarizes by source mixtures, analyzes with importance-parameterized ranking scores displayed as Tiles, and ranks methods stably; and (3) explicit application scenarios that fix data access, prior knowledge, and evaluation conditions so methods can be compared fairly.

What carries the argument

The probabilistic performance pipeline (performance as a probability measure on {tn, fp, fn, tp}, source-mixture summarization, and the two-parameter family of canonical ranking scores plotted as Tiles) that turns any well-specified random evaluation experiment into stable, preference-aware rankings.

Load-bearing premise

Every interesting video event can be turned, without leftover ambiguity about spatial or temporal boundaries, into a single two-class crisp classification whose four outcomes are fully defined by one random experiment that stays linear in both data sources and methods.

What would settle it

Apply the Tile ranking pipeline to a task whose event boundaries remain contested (for example, “how many people entered” under two different elementary-event definitions) and check whether the resulting orderings of methods reverse or become unstable across those definitions.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes a three-pillar framework for developing and comparing event-detection methods in videos: (i) a hierarchical, tag-based organization of event-monitoring video data spanning urban, natural, maritime, and underwater environments, including two new datasets (FSD from public IP cameras and synthetic SUC from CARLA); (ii) a probabilistic evaluation pipeline that casts detection as two-class crisp classification, averages performances via mixture-linear confusion matrices, and ranks methods with application-parameterized canonical ranking scores visualized as Tiles; and (iii) an explicit checklist of application-scenario characteristics (data, methods, evaluation, operational constraints) so that methods can be compared under stated conditions. The evaluation mathematics is grounded in the authors’ prior axiomatic work on ranking scores and Tiles; the dataset pillar unifies 36 public sources plus FSD/SUC under environment–modality–tag metadata.

Significance. If adopted, the framework would reduce the long-standing practice of ranking background-subtraction and related detectors with ad-hoc score averages and poorly specified tasks, and would give challenge organizers a reproducible way to report full confusion matrices and multi-criterion rankings. The Tile geometry and the linearity-based summarization/ranking results are carefully referenced to peer-reviewed foundations and are a genuine strength. FSD (≈153k annotated pairs) and SUC (≈187.5k frames with instance masks and weather metadata) are concrete, usable resources. The scenario checklist is a practical contribution for reproducibility and for high-risk AI documentation. The main value is organizational and methodological rather than a new detector or a large empirical leaderboard.

major comments (3)
  1. [§III–V (esp. §IV pipeline and new datasets)] The central claim is that the three pillars form a usable development framework, yet the manuscript never runs the full pipeline end-to-end on FSD or SUC (or on a multi-environment subset of EMVD). Section IV illustrates Tiles on CDnet and cites IWDD, but there is no Entity/Value Tile, mixture summarization, or scenario-conditioned ranking produced from the new data. Without at least one complete worked example, the claim that the framework enables rigorous development remains aspirational rather than demonstrated.
  2. [§IV-A1–A4] §IV largely restates the probabilistic performance model, ranking scores R_I, canonical scores parameterized by (a,b), and Tile flavors from the authors’ prior work [13,14,39]. The event-detection-specific content is mainly the three random-evaluation experiments and the hardware/real-time metrics. The manuscript should state explicitly what is new for video event detection versus what is imported, and should show where the linearity-in-method / linearity-in-source assumptions fail for common tasks (e.g., multi-object matching with non-unique associations, or clip-level events with soft temporal bounds).
  3. [§I, §IV-A1, §V-C] The paper correctly flags Bertrand’s paradox and ill-posed spatial/temporal bounds (§I), then requires the designer to fix a single linear random evaluation experiment. That is a scope condition, not a contradiction, but it is load-bearing: many clip-level and action-spotting tasks do not admit an unambiguous exhaustive negative class or a unique matching rule. §V should require every scenario to publish the exact random experiment (or matching protocol) and to state whether linearity holds; otherwise Tile rankings lose their formal guarantees on those tasks.
minor comments (6)
  1. [Abstract, §I] Abstract and §I claim that lack of large-scale data and rigorous evaluation “have biased” comparisons; this is plausible but unsupported by a concrete before/after or citation analysis. Soften or cite evidence.
  2. [§III-B, Table I, Fig. 2] Fig. 1 and Table I are useful; ensure every dataset listed in §III-B appears consistently in the table and timeline (Fig. 2), including Audio-Visual Vehicle and FSD/SUC.
  3. [§III-C] FSD and SUC are partly private “to enable fair evaluation.” State clearly what is public now, what will be released, and how third parties can reproduce the framework without the private split.
  4. [§IV-D, §IV-E, §V-D] Hardware metrics (§IV-D–E) are sensible but disconnected from the probabilistic pipeline; a short note on how MEM_Δ / PFR_Δ / delay interact with scenario constraints would help.
  5. [Title page, throughout] Minor typos and formatting: “Deli `ege”, “Pi ´erard”, “Staff, IEEE,” trailing commas in author list; “W ACV” spacing; arXiv id in header is future-dated (2607).
  6. [§IV-B (G2)] Guideline (G2) urges full confusion matrices; the paper itself does not release matrices for FSD/SUC. Align practice with the guideline or mark them as forthcoming.

Circularity Check

1 steps flagged

Mild self-citation load-bearing for the evaluation pillar only; datasets, scenarios, and overall three-pillar proposal remain independent of any constructional reduction.

specific steps
  1. self citation load bearing [Section II (Related Work) and Section IV-A (esp. IV-A1–A4)]
    "The second limitation has been addressed in the works of Piérard et al. [13, 14] which have been applied successfully in the International Contest on Illegal Waste Dumping Detection (IWSS 2026) in conjunction with WACV 2026 [7]. These works have led us to propose the performance evaluation tools described in Section IV. ... we follow the framework of Piérard et al. [13] ... Piérard et al. [13] introduced the first axiomatic framework for performance-based rankings ... the canonical ranking scores are defined in [14, 39]"

    One of the three main pillars (the ‘rigorous’ probabilistic evaluation pipeline, ranking scores, and Tile visualizations) is justified almost exclusively by citations to prior work whose author lists heavily overlap the present paper. The present text treats those results as established external foundations rather than re-deriving them, making the evaluation claims load-bearing on self-citation. The prior works themselves contain independent mathematical arguments, so the reduction is not total; the other two pillars (datasets, scenarios) do not depend on it.

full rationale

This is a framework/proposal paper, not a derivation of a numerical prediction or uniqueness theorem from first principles. The three claimed pillars (hierarchical tagged datasets including new FSD/SUC collections, a probabilistic evaluation pipeline with Tiles/ranking scores, and explicit application scenarios) do not reduce by construction to their own inputs. The evaluation pillar (Section IV) does rest load-bearingly on the authors’ own prior axiomatic ranking and Tile papers (Piérard et al. [13,14] and related), which share multiple co-authors and are invoked as the foundation for summarization, canonical ranking scores, Value/Entity Tiles, and stability arguments. Those citations supply independent mathematical content rather than fitted parameters or tautological redefinitions, so the circularity is limited to self-citation dependence for one pillar. No equation equates a claimed prediction to a fitted input; no uniqueness theorem is imported solely to forbid alternatives; Bertrand’s paradox is acknowledged as a scope condition rather than resolved circularly. Datasets and scenario language stand free of that chain. Score 2 reflects proportionate mild self-citation without forcing the central framework claim.

Axiom & Free-Parameter Ledger

2 free parameters · 3 axioms · 3 invented entities

The central claim is a methodological proposal rather than a numerical derivation; free parameters are therefore few and mostly application-chosen rather than fitted. The load-bearing axioms are the probabilistic modeling choices and the linearity assumptions required for the ranking theory. Invented entities are the concrete new datasets and the formalized notion of an application scenario.

free parameters (2)
  • mixture weights λ_s over data sources
    Chosen by the evaluator when summarizing performances; arbitrary non-negative weights summing to 1. Not fitted to any target score inside the paper, but free and application-dependent.
  • importance pair (a,b) on the Tile
    Relative importance of true positives vs. true negatives and of false negatives vs. false positives. Free parameters that encode user preference; the paper does not fit them to data.
axioms (3)
  • domain assumption Any event-detection task of interest admits an exhaustive two-class crisp classification formulation whose four outcomes are completely defined by a single random evaluation experiment.
    Stated in §I and used throughout §IV; if event boundaries remain ambiguous the probabilistic scores lose meaning.
  • domain assumption The chosen random evaluation experiment is linear with respect to both the data source and the evaluated method, so that mixtures of videos or of methods yield convex combinations of performances.
    Required for the summarization formula (Eq. 11) and the hybrid-method ranking guarantee (Eq. 14) in §IV-A2 and §IV-A4.
  • standard math Standard probability measure theory on the finite sample space Ω = {tn,fp,fn,tp} is an adequate model of performance.
    Background for the entire §IV-A pipeline; taken from the authors’ prior ranking theory papers.
invented entities (3)
  • Foreground Segmentation Dataset (FSD) no independent evidence
    purpose: Real-world multi-camera IP-camera collection with pixel masks and challenge tags for foreground segmentation evaluation.
    Newly collected; 153 k annotated pairs; partially private. No independent public release yet.
  • Synthetic Urban Crossroad (SUC) dataset no independent evidence
    purpose: CARLA-generated urban traffic sequences with instance masks and weather metadata for controlled evaluation.
    Newly generated; 187.5 k frames; partially private. Independent evidence limited to the description inside this paper.
  • Application scenario (formal checklist) no independent evidence
    purpose: Explicit declaration of data access, prior knowledge, causality, hardware limits, and evaluation protocol so that methods become comparable.
    Conceptual construct introduced in §V; no external validation that the checklist is complete or that it eliminates all ambiguity.

pith-pipeline@v1.1.0-grok45 · 36224 in / 3002 out tokens · 42999 ms · 2026-07-11T19:39:58.816451+00:00 · methodology

0 comments
read the original abstract

Event detection tasks in videos, the most important aspect of video surveillance, aim to detect events either at the pixel-level, frame-level, or clip-level. Plenty of methods intended for event detection in different environments, for various applications, and within different acquisition techniques were introduced. Naturally, the attempts were made as well to classify these algorithms in terms of detection of performance or in terms of real-time abilities. Nevertheless, the lack of a large-scale dataset as well as rigorous performance evaluation methods have biased such comparisons as well as the development of the methods. Given the diversity of existing approaches, we believe it is essential for researchers to position their work within such a rich landscape. Thus, we propose a rigorous framework for developing new methods in event detection for videos. Specifically, this framework is based on three main pillars: datasets, performance evaluation, and scenarios for deploying methods.

Figures

Figures reproduced from arXiv: 2607.04372 by Adrien Deli\`ege, Ana\"is Halin, Anastasia Zakharova, Anthony Cioppa, Antonio Greco, Bruno Vento, Carlo Sansone, Islam Osman, Kamil Jeziorek, Marc Van Droogenbroeck, Meghna Kapoor, Mohamed S. Shehata, Renaud Vandeghen, S\'ebastien Pi\'erard, Thierry Bouwmans, Tomasz Kryjak.

Figure 1
Figure 1. Figure 1: Categorization of the Event-Monitoring Video Dataset According to Environment and Modality [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Event Detection in Videos Datasets Timeline for Each Category: Urban Small-scale datasets, Urban Large-scale datasets, Maritime datasets, Underwater [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Representative samples from the Foreground Segmentation Dataset [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: An RGB frame with its corresponding instance mask (from the SUC [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: First RGB frame of each of the 25 video sequences of the SUC dataset. This illustrates the various view points, illumination, and weather conditions across the different sequences. the task to the equipment used and the complexity of the method. This section addresses several of these aspects. A. A Probabilistic Performance Evaluation Pipeline In this section, we present a complete and generic pipeline for… view at source ↗
Figure 6
Figure 6. Figure 6: Illustration of a complete pipeline for evaluating and comparing the performance of methods, applied to the task of background subtraction on CDnet. [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Two interpretative readings of the Tile: a map of application-specific [PITH_FULL_IMAGE:figures/full_fig_p010_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Entity Tile, showing the best method for each [PITH_FULL_IMAGE:figures/full_fig_p012_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

98 extracted references · 35 canonical work pages · 2 internal anchors

  1. [1]

    Changedetection.net: A new change detection benchmark dataset

    Nil Goyette, Pierre-Marc Jodoin, Fatih Porikli, Janusz Konrad, and Prakash Ishwar. “Changedetection.net: A new change detection benchmark dataset”. In:IEEE Int. Conf. Comput. Vis. Pattern Recognit. Work. (CVPRW). Providence, RI, USA: IEEE, June 2012, pp. 1–8.URL: https://doi.org/10.1109/CVPRW.2012.6238919

  2. [2]

    CD- net 2014: An Expanded Change Detection Benchmark Dataset

    Yi Wang, Pierre-Marc Jodoin, Fatih Porikli, Janusz Konrad, Yannick Benezeth, and Prakash Ishwar. “CD- net 2014: An Expanded Change Detection Benchmark Dataset”. In:IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work. (CVPRW). Columbus, OH, USA: Inst. Electr. Electron. Eng. (IEEE), June 2014, pp. 393–400. URL: https://doi.org/10.1109/CVPRW.2014.126

  3. [3]

    SoccerNet: A Scalable Dataset for Action Spotting in Soccer Videos

    Silvio Giancola, Mohieddine Amine, Tarek Dghaily, and Bernard Ghanem. “SoccerNet: A Scalable Dataset for Action Spotting in Soccer Videos”. In:IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work. (CVPRW). Salt Lake City, UT, USA: IEEE, June 2018, pp. 1792– 179210.URL: https : / / doi . org / 10 . 1109 / cvprw. 2018 . 00223

  4. [4]

    Smoke detection in video using convolutional neural networks and efficient spatio-temporal features

    Mahdi Hashemzadeh, Nacer Farajzadeh, and Milad Heydari. “Smoke detection in video using convolutional neural networks and efficient spatio-temporal features”. In:Appl. Soft Comput.128 (Oct. 2022), p. 109496.URL: https://doi.org/10.1016/j.asoc.2022.109496

  5. [5]

    VideoGasNet: Deep learning for natural gas methane leak classification using an infrared camera

    Jingfan Wang, Jingwei Ji, Arvind P. Ravikumar, Silvio Savarese, and Adam R. Brandt. “VideoGasNet: Deep learning for natural gas methane leak classification using an infrared camera”. In:Energy238 (Jan. 2022), p. 121516.URL: https://doi.org/10.1016/j.energy.2021. 121516

  6. [6]

    Perimeter Intrusion Detection by Video Surveillance: A Survey

    Devashish Lohani, Carlos Crispim-Junior, Quentin Barth´elemy, Sarah Bertrand, Lionel Robinault, and Laure Tougne Rodet. “Perimeter Intrusion Detection by Video Surveillance: A Survey”. In:Sensors22.9 (May 2022), pp. 1–28.URL: https://doi.org/10.3390/ s22093601

  7. [7]

    Illegal waste dump- ing detection

    Thierry Bouwmans, Antonio Greco, S ´ebastien Pi ´erard, Andrea Vincenzo Ricciardi, Carlo Sansone, Marc Van Droogenbroeck, and Bruno Vento. “Illegal waste dump- ing detection”. In:IEEE/CVF Winter Conf. Appl. Com- put. Vis. Work. (WACVW). Tucson, AZ, USA, Mar. 2026, pp. 539–548

  8. [8]

    Gauthier- Villars et fils, 1889

    Joseph Bertrand.Calcul des probabilit ´es. Gauthier- Villars et fils, 1889

  9. [9]

    A survey of approaches and trends in person re-identification

    Apurva Bedagkar-Gala and Shishir K. Shah. “A survey of approaches and trends in person re-identification”. In:Image Vis. Comput.32.4 (Apr. 2014), pp. 270–286. URL: https://doi.org/10.1016/j.imavis.2014.02.001

  10. [10]

    Wallflower: Principles and Practice of Background Maintenance

    Kentaro Toyama, John Krumm, Barry Brumitt, and Brian Meyers. “Wallflower: Principles and Practice of Background Maintenance”. In:IEEE Int. Conf. Comput. Vis. (ICCV). Kerkyra, Greece, Sept. 1999, pp. 255–261. URL: https://doi.org/10.1109/ICCV .1999.791228

  11. [11]

    Evaluation of background subtraction tech- niques for video surveillance

    Sebastian Brutzer, Benjamin Hoferlin, and Gunther Hei- demann. “Evaluation of background subtraction tech- niques for video surveillance”. In:IEEE Int. Conf. Comput. Vis. Pattern Recognit. (CVPR). Providence, RI, USA: IEEE, June 2011, pp. 1937–1944.URL: https : //doi.org/10.1109/CVPR.2011.5995508

  12. [12]

    A Benchmark Dataset for Outdoor Foreground/Background Extraction

    Antoine Vacavant, Thierry Chateau, Alexis Wilhelm, and Laurent Lequi `evre. “A Benchmark Dataset for Outdoor Foreground/Background Extraction”. In:Asian Conf. Comput. Vis. (ACCV). V ol. 7728. Lect. Notes 17 Comput. Sci. Springer Berl. Heidelb., Nov. 2012, pp. 291–300.URL: https : / / doi . org / 10 . 1007 / 978 - 3 - 642-37410-4 25

  13. [13]

    Foun- dations of the Theory of Performance-Based Rank- ing

    S ´ebastien Pi ´erard, Ana ¨ıs Halin, Anthony Cioppa, Adrien Deli `ege, and Marc Van Droogenbroeck. “Foun- dations of the Theory of Performance-Based Rank- ing”. In:IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR). Nashville, TN, USA: IEEE, June 2025, pp. 14293–14302.URL: https : / / doi . org / 10 . 1109 / CVPR52734.2025.01333

  14. [14]

    The Tile: A 2D Map of Ranking Scores for Two-Class Classification

    S ´ebastien Pi ´erard, Ana ¨ıs Halin, Anthony Cioppa, Adrien Deli `ege, and Marc Van Droogenbroeck. “The Tile: A 2D Map of Ranking Scores for Two-Class Classification”. In:arXivabs/2412.04309 (2024). arXiv: 2412.04309.URL: https://doi.org/10.48550/arXiv.2412. 04309

  15. [15]

    CARLA: An Open Urban Driving Simulator

    Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun. “CARLA: An Open Urban Driving Simulator”. In:Annu. Conf. Robot. Learn.V ol. 78. Proc. Mach. Learn. Res. Mountain View, CA, USA: ML Research Press, Nov. 2017, pp. 1–16. URL: https://proceedings.mlr.press/v78/dosovitskiy17a. html

  16. [16]

    Real-Time Semantic Background Subtrac- tion

    Anthony Cioppa, Marc Van Droogenbroeck, and Marc Braham. “Real-Time Semantic Background Subtrac- tion”. In:IEEE Int. Conf. Image Process. (ICIP). Abu Dhabi, United Arab. Emir.: IEEE, Oct. 2020, pp. 3214– 3218.URL: https://doi.org/10.1109/ICIP40778.2020. 9190838

  17. [17]

    BSUV-Net 2.0: Spatio-Temporal Data Augmen- tations for Video-AgnosticSupervised Background Sub- traction

    M. Ozan Tezcan, Prakash Ishwar, and Janusz Kon- rad. “BSUV-Net 2.0: Spatio-Temporal Data Augmen- tations for Video-AgnosticSupervised Background Sub- traction”. In:IEEE Access9 (2021), pp. 53849–53860. URL: https://doi.org/10.1109/ACCESS.2021.3071163

  18. [18]

    Compar- ative study of background subtraction algorithms

    Yannick Benezeth, Pierre-Marc Jodoin, Bruno. Emile, H´el`ene Laurent, and Christophe Rosenberger. “Compar- ative study of background subtraction algorithms”. In: J. Electron. Imaging19.3 (July 2010), pp. 1–12.URL: https://doi.org/10.1117/1.3456695

  19. [19]

    A Probabilistic In- terpretation of Precision, Recall and F-Score, with Im- plication for Evaluation

    Cyril Goutte and Eric Gaussier. “A Probabilistic In- terpretation of Precision, Recall and F-Score, with Im- plication for Evaluation”. In:Advances in Information Retrieval (Proceedings of ECIR). V ol. 3408. Lect. Notes Comput. Sci. Springer, Mar. 2005, pp. 345–359.URL: https://doi.org/10.1007/978-3-540-31865-1 25

  20. [20]

    Nouvelles recherches sur la distribution florale

    Paul Jaccard. “Nouvelles recherches sur la distribution florale”. In:Bull. De La Soci ´et´e Vaudoise Des Sci. Nat. 44.163 (1908), pp. 223–270

  21. [21]

    Cornelis Joost van Rijsbergen.Information Retrieval. second. London, Engl.: Butterworths, 1979

  22. [22]

    A Review of the F-Measure: Its History, Properties, Criticism, and Alternatives

    Peter Christen, David J. Hand, and Nishadi Kirielle. “A Review of the F-Measure: Its History, Properties, Criticism, and Alternatives”. In:ACM Comput. Surv. 56.3 (Oct. 2023), pp. 1–24.URL: https://doi.org/10. 1145/3606367

  23. [23]

    PToPI: A Comprehensive Review, Analysis, and Knowledge Representation of Binary Classification Performance Measures/Metrics

    G ¨urol Canbek, Tugba Taskaya Temizel, and Seref Sa- giroglu. “PToPI: A Comprehensive Review, Analysis, and Knowledge Representation of Binary Classification Performance Measures/Metrics”. In:SN Computer Sci- ence4.1 (Oct. 2022).URL: https://doi.org/10.1007/ s42979-022-01409-1

  24. [24]

    Classification assessment methods

    Alaa Tharwat. “Classification assessment methods”. In: Appl. Comput. Informatics17.1 (2021), pp. 168–192. URL: https://doi.org/10.1016/j.aci.2018.08.003

  25. [25]

    Evaluation: from precision, recall and F-measure to ROC, informedness, marked- ness and correlation

    David M. W. Powers. “Evaluation: from precision, recall and F-measure to ROC, informedness, marked- ness and correlation”. In:arXivabs/2010.16061 (2020). arXiv: 2010 . 16061.URL: https : / / doi . org / 10 . 48550 / arXiv.2010.16061

  26. [26]

    Performance Measures for Binary Clas- sification

    Daniel Berrar. “Performance Measures for Binary Clas- sification”. In:Encycl. Bioinform. Comput. Biology (2019), pp. 546–560.URL: https : / / doi . org / 10 . 1016 / B978-0-12-809633-8.20351-8

  27. [27]

    A systematic analysis of performance measures for classification tasks

    Marina Sokolova and Guy Lapalme. “A systematic analysis of performance measures for classification tasks”. In:Inf. Process. & Manag.45.4 (July 2009), pp. 427–437.URL: https://doi.org/10.1016/j.ipm.2009. 03.002

  28. [28]

    An Analysis of Performance Measures for Binary Classifiers

    Charles Parker. “An Analysis of Performance Measures for Binary Classifiers”. In:IEEE Int. Conf. Data Min. Vancouver, Can.: IEEE, Dec. 2011, pp. 517–526.URL: https://doi.org/10.1109/ICDM.2011.21

  29. [29]

    An experimental comparison of performance measures for classification

    C `esar Ferri, Jos ´e Hern´andez-Orallo, and Ramona Mod- roiu. “An experimental comparison of performance measures for classification”. In:Pattern Recognit. Lett. 30.1 (Jan. 2009), pp. 27–38.URL: https://doi.org/10. 1016/j.patrec.2008.08.010

  30. [30]

    Beweis der invarianz desn- dimensionalen gebiets

    Luitzen E. J. Brouwer. “Beweis der invarianz desn- dimensionalen gebiets”. In:Math. Ann.71.3 (Sept. 1911), pp. 305–313.URL: https : / / doi . org / 10 . 1007 / BF01456846

  31. [31]

    Zur Invarianz desn- dimensionalen Gebiets

    Luitzen E. J. Brouwer. “Zur Invarianz desn- dimensionalen Gebiets”. In:Math. Ann. 72.1 (Mar. 1912), pp. 55–56.URL: https : //doi.org/10.1007/BF01456889

  32. [32]

    Ad- vancing Precision, Recall, F-score, and Jaccard index: An approach for continuous, ratio-scale measurements

    Katarzyna Krasnodebska, Wojciech Goch, Johannes H. Uhl, Judith A. Verstegen, and Martino Pesaresi. “Ad- vancing Precision, Recall, F-score, and Jaccard index: An approach for continuous, ratio-scale measurements”. In:Environ. Model. & Softw.193 (Sept. 2025), pp. 1–9. URL: https://doi.org/10.1016/j.envsoft.2025.106614

  33. [33]

    Frequentist Probability and Frequentist Statistics

    Jerzy Neyman. “Frequentist Probability and Frequentist Statistics”. In:Synthese36.1 (1977), pp. 97–131.URL: https://www.jstor.org/stable/20115217

  34. [34]

    The Relationship Between Precision-Recall and ROC Curves

    Jesse Davis and Mark Goadrich. “The Relationship Between Precision-Recall and ROC Curves”. In:Int. Conf. Mach. Learn. (ICML). Pittsburgh, Pennsylvania: ML Res. Press, June 2006, pp. 233–240

  35. [35]

    Sum- marizing the performances of a background subtraction algorithm measured on several videos

    S ´ebastien Pi´erard and Marc Van Droogenbroeck. “Sum- marizing the performances of a background subtraction algorithm measured on several videos”. In:IEEE Int. Conf. Image Process. (ICIP). Abu Dhabi, United Arab. Emir., Oct. 2020, pp. 3234–3238.URL: https://doi.org/ 10.1109/ICIP40778.2020.9190865

  36. [36]

    Unachievable Region in Precision-Recall Space 18 and Its Effect on Empirical Evaluation

    Kendrick Boyd, Victor Costa, Jesse Davis, and David Page. “Unachievable Region in Precision-Recall Space 18 and Its Effect on Empirical Evaluation”. In:Int. Conf. Mach. Learn. (ICML). Edinburgh, UK, June 2012, pp. 639–646

  37. [37]

    Multi-domain performance analysis with scores tailored to user preferences

    S ´ebastien Pi ´erard, Adrien Deli `ege, and Marc Van Droogenbroeck. “Multi-domain performance analysis with scores tailored to user preferences”. In:arXiv abs/2512.08715 (2025). arXiv: 2512.08715.URL: https: //doi.org/10.48550/arXiv.2512.08715

  38. [38]

    Unpublished work, submitted to ESANN

    S ´ebastien Pi ´erard, Adrien Deli `ege, and Marc Van Droogenbroeck.Multi-domain performance analysis with scores tailored to user preferences. Unpublished work, submitted to ESANN. 2025

  39. [39]

    A Hitchhiker's Guide to Understanding Performances of Two-Class Classifiers

    Ana ¨ıs Halin, S ´ebastien Pi ´erard, Anthony Cioppa, and Marc Van Droogenbroeck. “A Hitchhiker’s Guide to Understanding Performances of Two-Class Classifiers”. In:arXivabs/2412.04377 (2024). arXiv: 2412.04377. URL: https://doi.org/10.48550/arXiv.2412.04377

  40. [40]

    A Methodology to Evaluate Strategies Predicting Rankings on Unseen Domains

    S ´ebastien Pi ´erard, Adrien Deli `ege, Ana ¨ıs Halin, and Marc Van Droogenbroeck. “A Methodology to Evaluate Strategies Predicting Rankings on Unseen Domains”. In:IEEE Int. Conf. Multimedia Expo Work. (ICMEW), Work. Big Surveill. Data Anal. Process. (big-surv). Nantes, France: IEEE, June 2025, pp. 1–6.URL: https: //doi.org/10.1109/ICMEW68306.2025.11152102

  41. [41]

    Model Cards for Model Reporting

    Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. “Model Cards for Model Reporting”. In:Proc. Conf. Fairness, Accountability, Transpar.Atlanta, GA, USA: ACM, Jan. 2019, pp. 220–229.URL: https://doi.org/10. 1145/3287560.3287596

  42. [42]

    Sorbetto: a Python library for producing classification tiles with different flavors to visualize, analyze, and compare performances in two- class classification problems

    S ´ebastien Pi´erard, Ana¨ıs Halin, Franc ¸ois Marelli, Simon Pernas, and J ´erˆome Pierre. “Sorbetto: a Python library for producing classification tiles with different flavors to visualize, analyze, and compare performances in two- class classification problems”. In:Zenodo(Sept. 2025). URL: https://doi.org/10.5281/zenodo.17591788

  43. [43]

    Tornado predictions

    John P. Finley. “Tornado predictions”. In:Am. Meteorol. J.1.3 (July 1884), pp. 85–88.URL: https://archive.org/ details/sim american-meteorological-journal 1884-07 1 3/page/84/mode/2up

  44. [44]

    Communications Through Limited Response Ques- tioning

    Edward M. Bennett, Renee Alpert, and A. C. Goldstein. “Communications Through Limited Response Ques- tioning”. In:Public Opin. Q.18.3 (1954), pp. 303–308. URL: https://doi.org/10.1086/266520

  45. [45]

    Reliability of Content Analysis: The Case of Nominal Scale Coding

    William A. Scott. “Reliability of Content Analysis: The Case of Nominal Scale Coding”. In:Public Opin. Q. 19.3 (1955), pp. 321–325.URL: https : / / doi . org / 10 . 1086/266577

  46. [46]

    A Coefficient of Agreement for Nom- inal Scales

    Jacob Cohen. “A Coefficient of Agreement for Nom- inal Scales”. In:Educ. Psychol. Meas.20.1 (Apr. 1960), pp. 37–46.URL: https : / / doi . org / 10 . 1177 / 001316446002000104

  47. [47]

    A Fallacy in the Use of Skill Scores

    Herbert S. Appleman. “A Fallacy in the Use of Skill Scores”. In:Bull. Am. Meteorol. Soc.41.2 (Feb. 1960), pp. 64–67.URL: https://doi.org/10.1175/1520- 0477- 41.2.64

  48. [48]

    Finley’s tornado predictions

    Grove Karl Gilbert. “Finley’s tornado predictions”. In: Am. Meteorol. J.1.5 (1884), pp. 166–172

  49. [49]

    The numerical measure of the success of predictions

    Charles S. Peirce. “The numerical measure of the success of predictions”. In:Science4.93 (Nov. 1884), pp. 453–454.URL: http://www.jstor.org/stable/1760565

  50. [50]

    On the association of attributes in statistics: with illustrations from the material of the childhood society, &c

    George Udny Yule. “On the association of attributes in statistics: with illustrations from the material of the childhood society, &c.” In:Philosophical Transactions of the Royal Society of London. Series A, Contain- ing Papers of a Mathematical or Physical Character 194.252-261 (1900), pp. 257–319.URL: https://doi.org/ 10.1098/rsta.1900.0019

  51. [51]

    Berechnung Des Erfolges Und Der G ¨ute Der Windst ¨arkevorhersagen Im Sturmwarnungsdienst

    Paul Heidke. “Berechnung Des Erfolges Und Der G ¨ute Der Windst ¨arkevorhersagen Im Sturmwarnungsdienst”. In:Geografiska Annaler8.4 (Dec. 1926), pp. 301–349. URL: https://doi.org/10.1080/20014422.1926.11881138

  52. [52]

    Rating weather forecasts

    H. Helm Clayton. “Rating weather forecasts”. In:Bull. Am. Meteorol. Soc.15.12 (Dec. 1934), pp. 279–283. URL: https://doi.org/10.1175/1520-0477-15.12.279

  53. [53]

    The theory of signal detectability

    Wesley W. Peterson, Theodore G. Birdsall, and W. C. Fox. “The theory of signal detectability”. In:Trans. IRE Prof. Group Inf. Theory4.4 (Sept. 1954), pp. 171–212. URL: https://doi.org/10.1109/TIT.1954.1057460

  54. [54]

    Measuring the accuracy of diagnostic systems

    John A. Swets. “Measuring the accuracy of diagnostic systems”. In:Science240 (1988), pp. 1285–1293

  55. [55]

    Area under the Precision-Recall Curve: Point Estimates and Confidence Intervals

    Kendrick Boyd, Kevin Eng, and David Page. “Area under the Precision-Recall Curve: Point Estimates and Confidence Intervals”. In:Eur. Conf. Mach. Learn. Princ. Pr. Knowl. Discov. Databases (ECML/PKDD). V ol. 8190. Lect. Notes Comput. Sci. Prague, Czech Repub.: Springer, Sept. 2013, pp. 451–466.URL: https: //doi.org/10.1007/978-3-642-40994-3 29

  56. [56]

    The Geometry of ROC Space: Under- standing Machine Learning Metrics through ROC Iso- metrics

    Peter A. Flach. “The Geometry of ROC Space: Under- standing Machine Learning Metrics through ROC Iso- metrics”. In:Int. Conf. Mach. Learn. (ICML). Washing- ton, DC, USA: ML Res. Press, Aug. 2003, pp. 194–201

  57. [57]

    A Novel Video Dataset for Change Detection Benchmarking

    Nil Goyette, Pierre-Marc Jodoin, Fatih Porikli, Janusz Konrad, and Prakash Ishwar. “A Novel Video Dataset for Change Detection Benchmarking”. In:IEEE Trans. Image Process.23.11 (Nov. 2014), pp. 4663–4679. URL: https://doi.org/10.1109/TIP.2014.2346013

  58. [58]

    An introduction to ROC analysis

    Tom Fawcett. “An introduction to ROC analysis”. In: Pattern Recognit. Lett.27.8 (June 2006), pp. 861–874. URL: https://doi.org/10.1016/j.patrec.2005.10.010

  59. [59]

    What Is the Optimal Ranking Score Between Precision and Recall? We Can Always Find It and It Is RarelyF 1

    S ´ebastien Pi ´erard, Adrien Deli `ege, and Marc Van Droogenbroeck. “What Is the Optimal Ranking Score Between Precision and Recall? We Can Always Find It and It Is RarelyF 1”. In:IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR). Denver, CO, USA: IEEE, June 2026

  60. [60]

    A unifying view on dataset shift in classification

    Jose G. Moreno-Torres, Troy Raeder, Roc ´ıo Alaiz- Rodr´ıguez, Nitesh V . Chawla, and Francisco Herrera. “A unifying view on dataset shift in classification”. In: Pattern Recognit.45.1 (Jan. 2012), pp. 521–530.URL: https://doi.org/10.1016/j.patcog.2011.06.019

  61. [61]

    Official Journal of the European Union, OJ L, 12.7.2024

    European Parliament and Council of the European Union.Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying 19 down harmonised rules on artificial intelligence (Artifi- cial Intelligence Act). Official Journal of the European Union, OJ L, 12.7.2024. 2024.URL: https://eur- lex. europa . eu / legal - content / EN / T...

  62. [62]

    Detecting moving shadows: algorithms and evaluation

    Andrea Prati, Ivana Mikic, Mohan M. Trivedi, and Rita Cucchiara. “Detecting moving shadows: algorithms and evaluation”. In:IEEE Trans. Pattern Anal. Mach. Intell. 25.7 (July 2003), pp. 918–923.URL: https://doi.org/10. 1109/TPAMI.2003.1206520

  63. [63]

    A Two-Stage Template Approach to Person Detection in Thermal Imagery

    James W. Davis and Mark A. Keck. “A Two-Stage Template Approach to Person Detection in Thermal Imagery”. In:IEEE Workshops on Applications of Com- puter Vision (WACV/MOTION). V ol. 1. Breckenridge, CO, USA: IEEE, Jan. 2005, pp. 364–369.URL: https: //doi.org/10.1109/ACVMOT.2005.14

  64. [64]

    Web site https:// vcipl-okstate.org/pbvs/bench/

    Roland Miezianko.IEEE OTCBVS WS Series Bench: Terravic Research Infrared Database. Web site https:// vcipl-okstate.org/pbvs/bench/. 2005.URL: https://vcipl- okstate.org/pbvs/bench/

  65. [65]

    ETISEO, performance eval- uation for video surveillance systems

    Anh-Tuan Nghiem, Franc ¸ois Bremond, Monique Thon- nat, and Val ´ery Valentin. “ETISEO, performance eval- uation for video surveillance systems”. In:IEEE Int. Conf. Adv. Video Signal Based Surveill. (AVSS). IEEE, Sept. 2007, pp. 476–481.URL: https://doi.org/10.1109/ A VSS.2007.4425357

  66. [66]

    Modeling, Clustering, and Segmenting Video with Mixtures of Dy- namic Textures

    Antoni B. Chan and Nuno Vasconcelos. “Modeling, Clustering, and Segmenting Video with Mixtures of Dy- namic Textures”. In:IEEE Trans. Pattern Anal. Mach. Intell.30.5 (May 2008), pp. 909–926.URL: https://doi. org/10.1109/TPAMI.2007.70738

  67. [67]

    Change Detection in Optical Aerial Images by a Multilayer Conditional Mixed Markov Model

    Csaba Benedek and Tam ´as Szir´anyi. “Change Detection in Optical Aerial Images by a Multilayer Conditional Mixed Markov Model”. In:IEEE Trans. Geosci. Remote Sens.47.10 (Oct. 2009), pp. 3416–3430.URL: https : //doi.org/10.1109/TGRS.2009.2022633

  68. [68]

    A large-scale benchmark dataset for event recognition in surveillance video

    Sangmin Oh, Anthony Hoogs, Amitha Perera, Naresh Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, J. K. Aggarwal, Hyungtae Lee, Larry Davis, Eran Swears, Xioyang Wang, Qiang Ji, Kishore Reddy, Mubarak Shah, Carl V ondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba, Bi Song, Anesco Fong, Amit Roy-Chowdhury, and Mita Desai. “A ...

  69. [69]

    Real time moving vehicle detection and reconstruction for improving classifica- tion

    Tao Wang and Zhigang Zhu. “Real time moving vehicle detection and reconstruction for improving classifica- tion”. In:IEEE Work. Appl. Comput. Vis. (WACV). Breckenridge, CO, USA: IEEE, Jan. 2012, pp. 497–502. URL: https://doi.org/10.1109/W ACV .2012.6163039

  70. [70]

    Background Subtraction Based on Color and Depth Using Active Sensors

    Enrique Fernandez-Sanchez, Javier Diaz, and Eduardo Ros. “Background Subtraction Based on Color and Depth Using Active Sensors”. In:Sensors13.7 (July 2013), pp. 1–21.URL: https : / / doi . org / 10 . 3390 / s130708895

  71. [71]

    Person Re-identification by Video Ranking

    Taiqing Wang, Shaogang Gong, Xiatian Zhu, and Shengjin Wang. “Person Re-identification by Video Ranking”. In:Eur. Conf. Comput. Vis. (ECCV). V ol. 8692. Lect. Notes Comput. Sci. Springer Int. Publ., 2014, pp. 688–703.URL: https://doi.org/10.1007/978- 3-319-10593-2 45

  72. [72]

    Towards Benchmarking Scene Background Initialization

    Lucia Maddalena and Alfredo Petrosino. “Towards Benchmarking Scene Background Initialization”. In: Int. Conf. Image Anal. Process. Work. (ICIAP Work.) V ol. 9281. Lect. Notes Comput. Sci. Springer Int. Publ., Sept. 2015, pp. 469–476.URL: https://doi.org/10.1007/ 978-3-319-23222-5 57

  73. [73]

    PETS 2016: Dataset and Challenge

    Luis Patino, Tom Cane, Alain Vallee, and James Ferryman. “PETS 2016: Dataset and Challenge”. In: IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work. (CVPRW). Las Vegas, NV , USA: IEEE, June 2016, pp. 1240–1247.URL: https://doi.org/10.1109/CVPRW. 2016.157

  74. [74]

    Weighted Low-rank Decomposi- tion for Robust Grayscale-Thermal Foreground Detec- tion

    Chenglong Li, Xiao Wang, Lei Zhang, Jin Tang, Hejun Wu, and Liang Lin. “Weighted Low-rank Decomposi- tion for Robust Grayscale-Thermal Foreground Detec- tion”. In:IEEE Trans. Circuits Syst. Video Technol.27.4 (Apr. 2017), pp. 725–738.URL: https://doi.org/10.1109/ TCSVT.2016.2556586

  75. [75]

    The SYNTHIA Dataset: A Large Collection of Synthetic Images for Se- mantic Segmentation of Urban Scenes

    German Ros, Laura Sellart, Joanna Materzynska, David Vazquez, and Antonio M. Lopez. “The SYNTHIA Dataset: A Large Collection of Synthetic Images for Se- mantic Segmentation of Urban Scenes”. In:IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR). Las Vegas, NV , USA: IEEE, June 2016, pp. 3234–3243.URL: https : //doi.org/10.1109/CVPR.2016.352

  76. [76]

    Extensive Benchmark and Survey of Modeling Methods for Scene Background Initial- ization

    Pierre-Marc Jodoin, Lucia Maddalena, Alfredo Pet- rosino, and Yi Wang. “Extensive Benchmark and Survey of Modeling Methods for Scene Background Initial- ization”. In:IEEE Trans. Image Process.26.11 (Nov. 2017), pp. 5244–5256.URL: https://doi.org/10.1109/ TIP.2017.2728181

  77. [77]

    Comparative Evaluation of Background Subtraction Algorithms in Remote Scene Videos Cap- tured by MWIR Sensors

    Guangle Yao, Tao Lei, Jiandan Zhong, Ping Jiang, and Wenwu Jia. “Comparative Evaluation of Background Subtraction Algorithms in Remote Scene Videos Cap- tured by MWIR Sensors”. In:Sensors17.9 (Aug. 2017), pp. 1–31.URL: https://doi.org/10.3390/s17091945

  78. [78]

    A Benchmarking Framework for Background Subtraction in RGBD Videos

    Massimo Camplani, Lucia Maddalena, Gabriel Moy ´a Alcover, Alfredo Petrosino, and Luis Salgado. “A Benchmarking Framework for Background Subtraction in RGBD Videos”. In:Int. Conf. Image Anal. Process. (ICIAP). V ol. 10590. Lect. Notes Comput. Sci. Springer Int. Publ., 2017, pp. 219–229.URL: https://doi.org/10. 1007/978-3-319-70742-6 21

  79. [79]

    MOTSynth: How Can Synthetic Data Help Pedestrian Detection and Tracking?

    Matteo Fabbri, Guillem Braso, Gianluca Maugeri, Or- cun Cetintas, Riccardo Gasparini, Aljosa Osep, Si- mone Calderara, Laura Leal-Taixe, and Rita Cucchiara. “MOTSynth: How Can Synthetic Data Help Pedestrian Detection and Tracking?” In:IEEE/CVF Int. Conf. Comput. Vis. (ICCV). Montr´eal, Can.: IEEE, Oct. 2021, pp. 10829–10839.URL: https : / / doi . org / 10...

  80. [80]

    AGVS: A New Change Detection Dataset for Airport Ground Video Surveillance

    Xiang Zhang, Chang Shu, Shuai Li, Celimuge Wu, and Zhi Liu. “AGVS: A New Change Detection Dataset for Airport Ground Video Surveillance”. In:IEEE Trans. Intell. Transp. Syst.23.11 (Nov. 2022), pp. 20588– 20600.URL: https://doi.org/10.1109/tits.2022.3184978

Showing first 80 references.