Pith. sign in

REVIEW 4 major objections 6 minor 32 references

A synthetic construction benchmark shows cartooning keeps worker-hazard detection usable; blur breaks it.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-08-01 22:37 UTC pith:KOWCE7XO

load-bearing objection SynthSite is a useful, honestly scoped synthetic benchmark for a rare relational hazard, but the headline privacy-utility ordering rests on 55 clips and low label agreement; the robust result is the worker-retention gap, not the F2 ranking. the 4 major comments →

arxiv 2607.16351 v1 pith:KOWCE7XO submitted 2026-07-17 cs.CV cs.AI

Privacy-Aware Synthetic Video Benchmarking and Relational Evaluation for Worker-Under-Suspended-Load Detection

classification cs.CV cs.AI
keywords worker under suspended loadsynthetic video benchmarkprivacy-aware generationrelational hazard detectionobfuscation evaluationconstruction safetyprivacy-utility trade-off
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper introduces SynthSite, a synthetic video benchmark for the rare relational hazard of a worker standing beneath a suspended load. It argues that whole-body obfuscations that preserve body structure, like cartooning, retain much of the geometric information needed for hazard recognition, while appearance-smoothing operations like blurring sharply reduce it. The paper also shows that a privacy transform's agreement with the raw-video pipeline is not the same as its agreement with human safety labels, so privacy evaluation must be human-grounded. If right, this shifts how privacy-utility trade-offs are measured for construction-safety analytics: preserve geometry, not appearance.

Core claim

The central empirical claim is that for worker-under-suspended-load recognition, structure-preserving whole-body obfuscations retain substantially more downstream utility than appearance-smoothing baselines, and that agreement with the raw pipeline is not identical to agreement with human hazard labels. Concretely, cartooning gives the strongest human-annotation F2 (0.767, vs 0.757 for raw), blur gives the weakest (0.671), and the dominant degradation is lost worker retention (41.4% for blur vs 88.9% for cartooning) rather than geometric drift (jitter <2.3% of bbox diagonal throughout).

What carries the argument

The carrying mechanism is a geometry-driven relational pipeline that (1) detects and tracks workers and loads, (2) infers a load's suspended state via a normalized clearance ratio between worker footpoints and load bottom edge, (3) defines a trapezoidal fall-zone proxy that widens toward the ground, and (4) aggregates per-frame hazard indicators over a one-second sliding window. This pipeline is evaluated under five candidate conditions; the obfuscation masks are shared, so differences come from the transform itself.

Load-bearing premise

The entire evaluation assumes that SynthSite's synthetic clips and their human labels are representative enough of real construction-site CCTV that the privacy-utility ranking transfers to operational footage.

What would settle it

A direct test would be to run the same five obfuscation conditions on a small set of real construction CCTV clips containing suspended loads; if blur outperforms cartooning on human-annotated hazard recognition there, the central ranking does not transfer.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If cartooning preserves hazard-recognition utility this well, construction-safety analytics can deploy privacy-preserving worker rendering without sacrificing safety detection.
  • Privacy evaluation for relational hazards should measure worker retention and localization stability, not just appearance suppression.
  • The distinction between raw-reference agreement and human-label agreement implies that privacy evaluations need human-annotated ground truth, not just pipeline consistency.
  • The benchmark offers a shareable testbed for rare suspended-load hazards that are dangerous to stage in real footage.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The finding that blur destroys worker retention suggests that other smoothing-based anonymization methods may harm downstream safety analytics in similar ways, a hypothesis testable on other relational tasks like worker-machinery proximity.
  • The paper's synthetic-only evaluation leaves open whether the ranking transfers to real CCTV; a natural extension is to repeat the five-condition evaluation on a small set of real suspended-load clips.
  • Because cartooning preserves silhouette while suppressing identity texture, it may also support finer-grained worker-centric tasks like PPE detection, though the paper notes this is unresolved.
  • The workflow of exporting textual intermediaries across a trust boundary could generalize to other rare-event benchmarks where raw footage cannot be released.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces SynthSite, a set of 55 synthetic construction-surveillance clips labeled for the relational hazard of a worker under a suspended load, together with a privacy-aware hybrid generation workflow that keeps raw footage inside a trusted environment and exports only textual intermediaries. The authors evaluate five whole-body worker obfuscation conditions (raw reference, Canny-edge, cartooning, pixelation, blur) using a lightweight open-vocabulary relational detection pipeline. Their central empirical claim is that structure-preserving obfuscations, especially cartooning, retain substantially more downstream hazard-recognition utility than appearance-smoothing baselines such as blur, and that agreement with the raw pipeline is not identical to agreement with human hazard labels. The paper is explicit that this is a controlled synthetic benchmark study and not a deployment-ready evaluation on real CCTV.

Significance. The paper addresses a genuinely under-resourced problem: relational construction hazards are rare and hard to capture, and privacy concerns limit data sharing. The workflow separating sensitive and non-sensitive generation is a useful practical contribution, and the controlled evaluation design — identical segmentation masks reused across obfuscations and identical pipeline thresholds — is a strength. The dataset and code release are also valuable. However, the headline quantitative claim rests on F2 differences computed from 55 clips and a human-label set with inter-annotator agreement of Cohen's κ = 0.308. The observed F2 differences are small (e.g., cartooning vs raw: 0.767 vs 0.757) and no uncertainty quantification or paired statistical test is provided. The large worker-retention gap (41.4% for blur vs 88.9% for cartooning) is likely robust, but the paper's central human-grounded conclusion needs additional statistical support before it can be accepted as stated.

major comments (4)
  1. [Section 3.4, Table 6] The human-annotation ground truth has Cohen's κ = 0.308, which conventionally indicates only fair agreement, and the headline F2 differences are within the expected noise at this sample size. On 55 clips with 27 unsafe events, Table 6 shows recall fixed at 0.852 for RAW, Canny-edge, Cartooning, and Pixelation, and precision values of 0.523 (RAW) and 0.548 (Cartooning). These imply about 23 true positives and roughly 44 vs 42 total positive predictions, so the best-vs-raw F2 gap corresponds to approximately two clip-level predictions. A one- or two-clip change can reorder the methods. No bootstrap confidence intervals, permutation tests, or McNemar tests are reported. This is load-bearing because the central claim that structure-preserving obfuscations 'retain substantially more downstream utility' depends on these F2 values. Please provide clip-level paired tests and uncertainty interval
  2. [Section 4.4, Table 2] The phrase 'substantially more downstream utility' is not supported for the human-annotation F2 results. The only large effect in Table 2 is worker retention (88.9% vs 41.4% for cartooning vs blur), which is persuasive. But the F2 ordering, cartooning 0.767, pixelation 0.762, raw 0.757, Canny-edge 0.742, blur 0.671, has adjacent differences smaller than 0.01, and the paper provides no evidence these are distinguishable from label noise. The claim that 'structure-preserving obfuscations best preserve relational hazard cues' should be restricted to what the data actually establish: a robust retention effect, plus a provisional F2 ranking that requires statistical validation.
  3. [Section 3.4, Appendix B] Reliability of the ground truth is central to any human-grounded evaluation. Cohen's κ = 0.308 is reported, but no per-category agreement, no confidence interval, and no analysis of whether adjudication changes label reliability are provided. Given that disagreements were resolved by a third annotator, the final labels are not necessarily better than the raw pairwise agreement suggests. Please report agreement on the unsafe/safe distinction separately, and consider a robustness analysis of the main F2 results under alternative adjudication rules or against the set of clips where both initial annotators agreed.
  4. [Section 4.5] The external-validity caveat about not testing on real CCTV is explicit and appropriate. However, the practical relevance of the privacy-utility ranking depends on transferability of the synthetic distribution to operational footage. Since the synthetic clips are generated from a small set of textual intermediaries and the LoRA internal-generation branch uses only 8 curated videos, I would like the authors to state more precisely what distribution the 55 clips are intended to represent, and to include at least a small qualitative sanity check against any available unlabeled real footage or a comparison of scene-level statistics. This is not necessarily a blocker for a benchmark paper, but it affects how strongly the conclusions can be generalized.
minor comments (6)
  1. [Abstract / Section 3.1] The dataset link is given, but the paper does not state the license or access conditions. Please add this detail.
  2. [Equation (4)] The notation d_raw is used for the diagonal of the raw bounding box. This is clear in context, but a one-sentence definition before the equation would help readers who jump directly to the metric.
  3. [Table 2] The table caption lists jitter metrics, but units are implied percentages. State explicitly that jitter is normalized by the raw bounding-box diagonal.
  4. [Figure 2 and Figure 3] The qualitative figures would benefit from a scale/zoom inset for the worker regions, since the workers are often small in wide surveillance frames, making it hard to visually verify the 'Canny-edge' and 'cartooned' differences.
  5. [Appendix D, Table 5] The parameter K_THRESHOLD is described as 'positive frames required to trigger UNSAFE.' Given that the window is 1 second (15 frames), K=10 corresponds to 67% occupancy, but the text in D.5 says 'at least 10 positive frames within the window.' Please clarify whether the latch is triggered immediately when the window contains 10 positives or only after a full 15-frame window.
  6. [Section 4.1] The statement that 'pixelation and blur serve as standard visual anonymization baselines' cites [10, 28]. It would also be useful to cite the specific bookkeeping needed for fair comparisons across obfuscations, such as the white contour added for pixelation and blur; this is described in Appendix C but not flagged in the main text.

Circularity Check

0 steps flagged

No circularity: the paper reports an empirical benchmark comparison with independent human labels; no fitted parameter is renamed as a prediction and no load-bearing self-citation is used.

full rationale

The paper's derivation chain is an empirical evaluation rather than a derivation from first principles. Section 4.3 defines raw-reference stability and agreement-with-human-annotations metrics, and Table 2 reports measured values. No equation defines an output in terms of the target claim: Eqs. (1)-(4) are geometric operationalizations (proximity filter, clearance ratio, hazard indicator, jitter) applied identically across all privacy conditions. The human-F2 values are produced by a fixed-threshold detection pipeline and compared with independent rubric-based annotations (Section 3.4), so human labels are not an input used to fit the pipeline. The obfuscation transformations are applied to shared person masks (Appendix C), making cross-condition differences measured outcomes rather than definitional consequences. The paper's central distinction between raw-reference agreement and human-label agreement is argued from the observed Table 2 values, not imposed by construction. There are no self-citations of the authors' prior results, no imported uniqueness theorem, and no ansatz smuggled in via citation. The acknowledged limitations (Section 4.5: no real CCTV; Appendix F.1: small precision differences between cartooning and raw) are external-validity and statistical-power concerns, which are correctness risks rather than circularity. Accordingly, no circular step is identified.

Axiom & Free-Parameter Ledger

9 free parameters · 4 axioms · 0 invented entities

No new physical entities are postulated. The paper's central claim depends on hand-set pipeline thresholds (Table 5), a synthetic-to-real representativeness assumption, and the reliability of modest-agreement human labels. These are the main things a reader is asked to accept without independent evidence.

free parameters (9)
  • SUSPENDED_CLEARANCE_RATIO = 0.12
    Hand-set threshold in Table 5; controls when a load is declared suspended and thereby affects every hazard decision.
  • NEARBY_WORKER_X_EXPANSION = 1.25
    Horizontal search radius (x load width) for nearby workers; hand-set and directly affects suspension inference.
  • K_THRESHOLD = 10 positive frames per 1s window
    Temporal decision threshold; hand-set to be tolerant, directly sets the clip-level UNSAFE output.
  • MAX_NEGATIVE_GAP = 2
    Temporal infilling tolerance; affects the persistence measure in Section 4.2.
  • V_DRIFT_MIN = 0.4 m/s
    Minimum lateral velocity floor for fall-zone width; hand-set in Appendix D.
  • ASSUMED_WORKER_HEIGHT_M = 1.75 m
    Scale normalization constant for pixel-to-meter conversion; errors shift fall-zone dimensions.
  • VERTICAL_GATE_PIXELS = 40
    Vertical compatibility gate; hand-set and suppresses or exposes worker-in-zone decisions.
  • YOLO_CONF = 0.15
    Detection confidence threshold; chosen in a pilot, affects worker and load retention numbers.
  • Obfuscation parameters = Canny(100,200), cartoon K=4, blur(51,51), pixelation grid
    One parameterization per obfuscation family; the paper acknowledges only one parameterization is tested (Section 4.5).
axioms (4)
  • domain assumption Human labels from two-annotator review with Cohen's kappa 0.308 are a reliable ground truth for the benchmark
    Section 3.4; low kappa indicates a nontrivial fraction of labels were disputed. Adjudication helps, but the 'human annotations' target remains noisy.
  • domain assumption SynthSite's synthetic clips are representative of operational suspended-load construction CCTV
    Section 4.5 explicitly states no real CCTV evaluation was performed. If the synthetic distribution differs in geometry, scale, or appearance, the privacy-utility ranking may not transfer.
  • domain assumption Worker footpoint (bbox bottom-center) and load bottom edge encode physically meaningful suspension geometry
    Appendix D.2; the monocular 2D proxy has no camera calibration, so depth-ambiguous scenes can be misread.
  • domain assumption Textual intermediaries crossing the trust boundary do not leak identity or site-identifying information
    Section 3.2 states this is an operational, not formal, privacy guarantee. Any leakage would weaken the privacy-aware workflow claim.

pith-pipeline@v1.3.0-alltime-deepseek · 14787 in / 12533 out tokens · 81164 ms · 2026-08-01T22:37:13.581974+00:00 · methodology

0 comments
read the original abstract

Publicly shareable construction-video benchmarks remain scarce, especially for safety-critical hazards that are rare, dangerous to stage, and difficult to release. We study worker under suspended load, a relational hazard that depends on worker-load geometry and temporal persistence rather than object detection alone. We introduce SynthSite, a focused synthetic video benchmark of 55 clips spanning varied load configurations, viewpoints, clutter, occlusions, and surveillance conditions, together with a privacy-aware hybrid generation workflow that supports both publicly shareable benchmark creation and privacy-constrained synthetic video generation. We then ask whether worker appearance can be suppressed without undermining downstream hazard recognition. Under five whole-body privacy conditions, we evaluate worker and load retention, localization stability, and clip-level hazard recognition. We find that structure-preserving obfuscations retain substantially more downstream utility than appearance-smoothing baselines, and that preserving a raw visual reference alone does not guarantee the strongest agreement with human hazard labels. These findings suggest that privacy evaluation for construction safety analytics should assess not only appearance suppression, but also preservation of the geometric cues required for hazard reasoning. Our dataset and code are available at https://huggingface.co/datasets/govtech/SynthSite .

Figures

Figures reproduced from arXiv: 2607.16351 by Alejandro Seif, Anshu Singh.

Figure 1
Figure 1. Figure 1: Privacy-aware hybrid workflow and representative SynthSite clips. Raw construction footage remains inside the sensitive environment, where a local vision–language model (Qwen2.5-VL) generates approved textual intermediaries. These intermediaries cross into the non-sensitive environment, where Nano Banana 2 produces synthetic seed images and Veo 3.1 Fast generates candidate publicly shareable synthetic vide… view at source ↗
Figure 2
Figure 2. Figure 2: Relational privacy evaluation under worker obfuscation. The same worker-under-suspended-load scene is shown under five evaluation conditions: a raw-video reference and four whole-body obfuscation settings: Canny-edge, cartooned, pixelated, and blurred worker representations. oblique, or occluded in construction surveillance. In such settings, identity-related cues may still be conveyed through soft biometr… view at source ↗
Figure 3
Figure 3. Figure 3: Qualitative examples from relational privacy evaluation under whole-body worker obfuscation. Each row shows the same scene under five conditions: raw, Canny-edge, cartooned, pixelated, and blurred worker representations. Orange overlays denote the estimated fall-zone proxy; red boxes indicate unsafe worker–load detections, and yellow boxes indicate safe workers. Consistent with the quantitative results, st… view at source ↗
Figure 4
Figure 4. Figure 4: Custom annotation platform used for SynthSite labeling. Annotators reviewed each clip together with an inline event￾specific rubric defining unsafe and safe cases, hard negatives, and exclusion criteria for ambiguous or low-quality clips. Although the platform supported multiple construction-safety event types, this paper uses only the worker-under-suspended-load rubric. In the interface, True Positive cor… view at source ↗
Figure 5
Figure 5. Figure 5: Representative failure cases of the relational hazard pipeline. From left to right: (FN) hook not visible, so the suspended load is not detected; (FN) load mislocalization, where an incorrect bounding box suppresses suspension inference by placing the load bottom too close to nearby workers; (FP) entire-crane-as-load confusion, which produces an inflated fall-zone proxy; and (FP) depth ambiguity, where mon… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

32 extracted references · 6 linked inside Pith

  1. [1]

    Evaluation of human visual privacy protection: A three-dimensional framework and benchmark dataset, 2025

    Sara Abdulaziz, Giacomo D’Amicantonio, and Egor Bon- darev. Evaluation of human visual privacy protection: A three-dimensional framework and benchmark dataset, 2025

  2. [2]

    Dataset and benchmark for detecting moving objects in construction sites.Automation in Con- struction, 122:103482, 2021

    Xuehui An, Li Zhou, Zuguang Liu, Chengzhi Wang, Pengfei Li, and Zhiwei Li. Dataset and benchmark for detecting moving objects in construction sites.Automation in Con- struction, 122:103482, 2021

  3. [3]

    SynthGuard: Redefining synthetic data generation with a scalable and privacy-preserving work- flow framework

    Eduardo Brito, Mahmoud Shoush, Kristian Tamm, Paula Etti, and Liina Kamm. SynthGuard: Redefining synthetic data generation with a scalable and privacy-preserving work- flow framework. InAvailability, Reliability and Security – ARES 2025, pages 193–211. Springer, 2025

  4. [4]

    Extracting training data from diffu- sion models

    Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagiel- ski, Vikash Sehwag, Florian Tramer, Borja Balle, Daphne Ip- polito, and Eric Wallace. Extracting training data from diffu- sion models. In32nd USENIX Security Symposium (USENIX Security 23), pages 5253–5270, 2023

  5. [5]

    StegaV AR: Privacy- preserving video action recognition via steganographic do- main analysis.arXiv preprint arXiv:2512.12586, 2025

    Lixin Chen, Chaomeng Chen, et al. StegaV AR: Privacy- preserving video action recognition via steganographic do- main analysis.arXiv preprint arXiv:2512.12586, 2025

  6. [6]

    YOLO-World: Real-time open-vocabulary object detection

    Tianheng Cheng, Lin Song, Yixiao Ge, Wenyu Liu, Xing- gang Wang, and Ying Shan. YOLO-World: Real-time open-vocabulary object detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16901–16911, 2024

  7. [7]

    Eugene Yan Tao Chian, Yang Miang Goh, Jing Tian, and Brian H.W. Guo. Dynamic identification of crane load fall zone: A computer vision approach.Safety Science, 156: 105904, 2022

  8. [8]

    Privacy-preserving visual anal- ysis: Training video obfuscation models without sensitive labels.Applied Intelligence, 54:6041–6052, 2024

    Sander De Coninck, Wei-Cheng Wang, Sam Leroux, Steven Bohez, and Pieter Simoens. Privacy-preserving visual anal- ysis: Training video obfuscation models without sensitive labels.Applied Intelligence, 54:6041–6052, 2024

  9. [9]

    SODA: A large-scale open site object detection dataset for deep learning in construction.Automation in Construc- tion, 142:104499, 2022

    Rui Duan, Hui Deng, Mao Tian, Yichuan Deng, and Jiarui Lin. SODA: A large-scale open site object detection dataset for deep learning in construction.Automation in Construc- tion, 142:104499, 2022

  10. [10]

    Pri- vacy protection vs

    ´Ad´am Erd´elyi, Thomas Winkler, and Bernhard Rinner. Pri- vacy protection vs. utility in visual data: An objective evalu- ation framework.Multimedia Tools and Applications, 77(2): 2285–2313, 2018

  11. [11]

    Latent anonymization for privacy-preserving video understanding.arXiv preprint arXiv:2511.08666, 2025

    Joseph Fioresi, Ishan Rajendrakumar Dave, Chen Chen, and Mubarak Shah. Latent anonymization for privacy-preserving video understanding.arXiv preprint arXiv:2511.08666, 2025

  12. [12]

    Nano banana 2 generative API

    Google. Nano banana 2 generative API. Proprietary Gener- ative AI API, 2026

  13. [13]

    Video models are zero-shot learners and reasoners.arXiv preprint arXiv:2509.20328, 2025

    Google DeepMind. Video models are zero-shot learners and reasoners.arXiv preprint arXiv:2509.20328, 2025

  14. [14]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models.arXiv preprint arXiv:2106.09685, 2021

  15. [15]

    DeepPrivacy2: To- wards realistic full-body anonymization

    H ˚akon Hukkel ˚as and Frank Lindseth. DeepPrivacy2: To- wards realistic full-body anonymization. InProceedings of the IEEE/CVF Winter Conference on Applications of Com- puter Vision (WACV), pages 1329–1338, 2023

  16. [16]

    A scoping review of privacy and util- ity metrics in medical synthetic data.npj Digital Medicine, 8(1):60, 2025

    Bayrem Kaabachi, J ´er´emie Despraz, Thierry Meurers, Karen Otte, Mehmed Halilovic, Bogdan Kulynych, Fabian Prasser, and Jean Louis Raisaro. A scoping review of privacy and util- ity metrics in medical synthetic data.npj Digital Medicine, 8(1):60, 2025

  17. [17]

    Workplace safety & health report, 2024.https : / / www

    Ministry of Manpower, Singapore. Workplace safety & health report, 2024.https : / / www . mom . gov . sg/- /media/mom/documents/safety- health/ reports - stats / wsh - national - statistics / wsh-national-stats-2024.pdf, 2025. Accessed: 2026-03-11

  18. [18]

    1926.1425 – keeping clear of the load.https : / / www

    Occupational Safety and Health Administration. 1926.1425 – keeping clear of the load.https : / / www . osha . gov / laws - regs / regulations / standardnumber / 1926 / 1926 . 1425, 2025. Ac- cessed: 2026-03-11

  19. [19]

    AI-Toolkit: The ultimate training toolkit for fine- tuning diffusion models.https : / / github

    Ostris. AI-Toolkit: The ultimate training toolkit for fine- tuning diffusion models.https : / / github . com / ostris/ai-toolkit, 2024. Accessed: 2026-03-12

  20. [20]

    En- hancing tower crane safety: A computer vision and deep learning approach.Engineering Proceedings, 53(1):38, 2023

    Parham Pazari, Nasim Didehvar, and Amin Alvanchi. En- hancing tower crane safety: A computer vision and deep learning approach.Engineering Proceedings, 53(1):38, 2023

  21. [21]

    Qwen2.5-VL technical report.arXiv preprint arXiv:2502.13923, 2025

    Qwen Team. Qwen2.5-VL technical report.arXiv preprint arXiv:2502.13923, 2025

  22. [22]

    EchoNet-Synthetic: Privacy-preserving video generation for safe medical data sharing

    Hadrien Reynaud, Qingjie Meng, Mischa Dombrowski, Ari- jit Ghosh, Thomas Day, Alberto Gomez, Paul Leeson, and Bernhard Kainz. EchoNet-Synthetic: Privacy-preserving video generation for safe medical data sharing. InMedi- cal Image Computing and Computer Assisted Intervention – MICCAI 2024, pages 285–295. Springer, 2024

  23. [23]

    Shrigandhi and Sachin R

    Meenakshi N. Shrigandhi and Sachin R. Gengaje. CSOD-24: Construction site object detection dataset for safety monitor- ing at construction site using deep learning.Journal of Inno- vative Image Processing, 7(1):182–206, 2025

  24. [24]

    Synthetic data–anonymisation groundhog day

    Theresa Stadler, Bristena Oprisanu, and Carmela Tron- coso. Synthetic data–anonymisation groundhog day. In31st USENIX Security Symposium (USENIX Security 22), pages 1451–1468, 2022

  25. [25]

    Wan: Open and advanced large-scale video gen- erative models.arXiv preprint arXiv:2503.20314, 2025

    Wan Team. Wan: Open and advanced large-scale video gen- erative models.arXiv preprint arXiv:2503.20314, 2025

  26. [26]

    Synthesize privacy-preserving high-resolution images via private textual intermediaries

    Haoxiang Wang, Zinan Lin, Da Yu, and Huishuai Zhang. Synthesize privacy-preserving high-resolution images via private textual intermediaries. InAdvances in Neural Infor- mation Processing Systems, 2025

  27. [27]

    Struck- by-falling-objects workplace fatal injuries in 2h2024

    Workplace Safety and Health Council, Singapore. Struck- by-falling-objects workplace fatal injuries in 2h2024. https : / / www . tal . sg / wshc / resources / newsletters / wsh - advisory / struck - by - falling - objects - workplace - fatal - injuries-in-2h2024, 2025. Accessed: 2026-03-11

  28. [28]

    A privacy-preserving action recognition system using action-specific edge detection

    Zhenyu Wu, Zhangyang Wang, Zhaowen Wang, and Hailin Jin. A privacy-preserving action recognition system using action-specific edge detection. InIEEE International Con- ference on Image Processing (ICIP), 2018

  29. [29]

    Development of a digital twin-based simula- tion system and a novel synthetic video dataset for enhancing computer vision in construction site safety

    Zhengyu Wu, Yuxiang Feng, Yiannis Demiris, and Panagio- tis Angeloudis. Development of a digital twin-based simula- tion system and a novel synthetic video dataset for enhancing computer vision in construction site safety. InProceedings of the European Conference on Computing in Construction, 2024

  30. [30]

    Depth anything: Unleashing the power of large-scale unlabeled data

    Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10371–10381, 2024

  31. [31]

    ByteTrack: Multi-object tracking by associating ev- ery detection box

    Yifu Zhang, Peize Sun, Yi Jiang, Dongdong Yu, Fucheng Weng, Zehuan Yuan, Ping Luo, Wenyu Liu, and Xinggang Wang. ByteTrack: Multi-object tracking by associating ev- ery detection box. InEuropean Conference on Computer Vision (ECCV), pages 1–21. Springer, 2022

  32. [32]

    Hauptmann

    Liang Zheng, Yi Yang, and Alexander G. Hauptmann. Per- son re-identification: Past, present and future.arXiv preprint arXiv:1610.02984, 2016. Privacy-Aware Synthetic Video Benchmarking and Relational Evaluation for Worker-Under-Suspended-Load Detection Supplementary Material Table 3. LoRA fine-tuning settings for Wan2.2-I2V-A14B trained on 8 curated publi...

This paper was first reviewed by deepseek-v4-flash on August 1, 2026.