Pith. sign in

REVIEW 4 major objections 7 minor 85 references

ErgoChat: a Visual Query System for the Ergonomic Risk Assessment of Construction Workers

T0 review · 4 major / 7 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Fine-tuning a vision-language model on a small expert-annotated set of construction photos lets it answer postural-risk queries at 96.5% accuracy and write captions that expert raters prefer over a generic model's captions.

desk verdict The dataset and the captioning gains are real; the VQA headline is noise, and the shared-source test set needs fixing before the superiority claim is robust. read the letter →

arxiv 2412.19954 v1 pith:OOJKYOU7 submitted 2024-12-27 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords generativeartificialintelligencevision-languagemodellargelanguageergonomicriskassessmentconstructionsafetyvisualquestionansweringimagecaptioningREBA
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to show that a vision-language model fine-tuned on a small, domain-specific set of image-text pairs can identify postural ergonomic risks of construction workers and describe them in human-like language. The resulting system, ErgoChat, answers yes/no questions about whether a worker is exposed to postural risk with 96.5% accuracy on the authors' 200-image test set. It also generates image captions that the nine automatic metrics and a panel of 50 ergonomics-knowledgeable raters judge as better than captions from the same architecture trained only on generic data. The authors contribute a new 1,900-pair dataset, annotated with REBA-based risk labels, to support training and testing of such systems. If the finding generalizes, it gives safety personnel a way to query site photos interactively and get readable, automated risk assessments without attaching sensors to workers.

What carries the argument

The engine is a three-part vision-language design: a frozen vision transformer turns an image into a sequence of visual tokens; a linear projection layer groups and maps those tokens into the embedding space of a 7-billion-parameter autoregressive language model; and the language model generates text from the combined visual and textual tokens. Two task-identifier tokens, '[vqa]' and '[caption]', are prepended to prompts so the same network can switch between answering a question and writing a free-form description. The operation that carries the argument is fine-tuning this pretrained network on the authors' REBA-annotated construction dataset, which teaches the model to map postural features of construction work onto ergonomic-risk language; the ground-truth captions were written by ergonomics experts, so the fine-tuning signal is expert knowledge rather than generic alt-text.

What would settle it

Assemble a test set of construction-site photos from sites, contractors, camera angles, lighting conditions, and worker attire never seen in the fine-tuning data, with near-duplicates automatically removed, and recompute the VQA accuracy and the nine caption metrics; if the fine-tuned model's advantage over the generic model collapses to near zero, the reported gains come from training-set overlap rather than generalizable ergonomic understanding.

Watch

Extended reading notes

Core claim

The central claim is that domain-specific fine-tuning, not a new architecture, is what makes a vision-language model useful for ergonomic risk assessment. Starting from a model that already has broad visual knowledge from generic image-text pretraining, the authors fine-tune it on 1,700 ergonomics-specific image-text pairs in which each image is labeled via the REBA observation method, with 'exposed' meaning risk at medium level or higher, and captions describe the worker's actions and postural hazards. On the held-out 200-image test set, the fine-tuned ErgoChat answers the fixed question 'Is the worker exposed to postural ergonomic risks?' correctly 96.5% of the time, versus 95% for the un-fine-tuned model, and its generated captions align with the ground-truth risk label 86% of the time by perplexity, versus 63.5% before fine-tuning. Across the nine caption metrics, the fine-tuned model improves over the generic model on the large majority of test images, and 50 expert questionnaire respondents selected the fine-tuned model's caption as more accurate 84.4% of the time, rating it on average 69.7% more accurate than the alternative.

Load-bearing premise

The load-bearing assumption is that the 200 test images are a fair, non-overlapping sample of construction-worker photos; they were collected from the same online source and annotated with the same protocol as the 1,700 fine-tuning images, and the paper does not report any duplicate or near-duplicate removal between partitions.

Editorial extensions

If this is right

  • A safety inspector could upload a single site photo and ask whether a worker is exposed to postural ergonomic risk, receiving an answer with approximately 96.5% accuracy on similarly sourced photos.
  • The same model could generate a readable narrative of a worker's risky postures, which could be used to draft ergonomic injury reports or to train safety personnel who lack ergonomics expertise.
  • The released 1,900-pair dataset gives other researchers a shared resource for vision-language ergonomic risk assessment, replacing the current reliance on ad hoc or small-scale data.
  • If the result holds, generic pretraining plus a small expert-labeled domain set is a viable recipe for specialized visual-safety tasks, potentially reducing the need for large in-domain data collection.
  • The modest VQA gain (95% to 96.5%) suggests that for simple yes/no risk questions the generic model is already competent; the real value of fine-tuning is in generating accurate explanatory captions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not report removing near-duplicate images between its fine-tuning and test partitions, even though both come from the same online source; the 96.5% figure is therefore best read as an upper bound on performance, and a cross-site held-out test would separate memorization from generalization.
  • The same recipe—take a general vision-language model, build a small REBA-labeled image-text set, fine-tune—should transfer to other postural risk domains such as warehousing, healthcare, or agriculture, where observational ergonomics is still manual.
  • Because the authors note that captions sometimes misdescribe workers' actions, the binding constraint appears to be visual perception rather than language generation; improving the visual encoder or input resolution may pay off more than adding caption data.
  • A direct test of prompt sensitivity is missing: all 200 test questions used one fixed prompt, so measuring accuracy under rephrased questions would reveal whether the model is answering from the image or from prompt cues.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper presents ErgoChat, a vision-language system for postural ergonomic risk assessment of construction workers. The system is built on MiniGPT-v2 (EVA ViT encoder, linear projection, LLaMA2-7B) and is fine-tuned on a newly curated dataset of 1,900 image-text pairs annotated with REBA-based ergonomic risk labels and descriptions. The authors evaluate the fine-tuned model against the same architecture without fine-tuning on a 200-sample test set, using VQA accuracy, perplexity-based risk identification, nine image-captioning metrics, and a forced-choice human evaluation by 50 questionnaire respondents. They report that ErgoChat achieves 96.5% VQA accuracy versus 95% for the baseline, large improvements in captioning metrics (e.g., ROUGE_r 0.15 to 0.40, METEOR 0.11 to 0.36), and an 84.4% expert preference rate for the fine-tuned descriptions.

Significance. If the reported results hold, the paper demonstrates that fine-tuning a multimodal large language model on domain-specific REBA-annotated data can produce an interactive, natural-language ergonomic risk assessment tool, an application that is currently underexplored. The proposed dataset, task tokens for VQA/captioning, and the comparison against a same-architecture baseline are useful methodological contributions, and the authors' plan to release code and annotations would support reproducibility. However, the strength of the central claim is currently limited by the small single-source test set, the absence of significance testing, and an underspecified perplexity-based evaluation. The idea is promising and the captioning improvements are substantial, but the evidence as presented is not yet robust enough to fully support the abstract's superiority claim.

major comments (4)
  1. [Section 4.1, Table 6] The 1.5-point VQA accuracy difference (96.5% vs 95.0%) is computed on 200 test samples and is reported without confidence intervals or a significance test. The 95% Wilson interval for 96.5% on n=200 is approximately [93.9%, 99.0%], which includes 95%, and a paired McNemar test would not reject equality at conventional levels. The abstract's headline claim that the VQA functionality delivers 96.5% accuracy and surpasses the baseline therefore needs either supporting statistics (e.g., confidence intervals, McNemar test) or a more modest framing that separates the model's absolute accuracy from the evidence of improvement over the baseline.
  2. [Section 3.2] The fine-tuning and test partitions are drawn from the same online source ('construction works' searches) and are annotated with the same REBA-based protocol, and no duplicate or near-duplicate image removal is reported. Because the test set is part of the same distribution on which the model is fine-tuned, the reported VQA and IC gains could reflect memorization of recurring scenes rather than generalizable ergonomic reasoning. The authors should report a duplicate/near-duplicate check (e.g., embedding-similarity screening) and ideally validate on an independently collected set; at minimum, the shared-source nature of the test set should be explicitly acknowledged as a threat to external validity in the limitations discussion in Section 5.
  3. [Sections 3.4.2 and 4.1] The use of perplexity as an ergonomic-risk classifier is not specified. Perplexity is a sequence-level likelihood, so converting it into the binary 'correct identification' rates in Table 6 requires a decision rule or threshold, which is not given. The assertion that a lower perplexity score implies that the generated description 'concludes the worker is exposed to ergonomic risk' is not self-evident and should be justified mechanically; without this, the 86% vs 63.5% perplexity result is uninterpretable as evidence of improved risk identification.
  4. [Section 3.5, Table 9] The human evaluation is labeled as 'assessments from human experts,' but the participant demographics in Section 4.2 show that only 10% self-identify as experts in ergonomic knowledge, while 46% are novices or have only fundamental awareness. The paper does not report inter-rater agreement or an expertise-stratified analysis, so it is unclear whether the 84.4% preference rate reflects ergonomic accuracy or generic descriptive quality. The authors should either reanalyze the forced-choice responses by self-reported expertise level or soften the 'expert' characterization in the abstract and conclusion.
minor comments (7)
  1. [Eq. (2), Section 3.4.2] The formula for ROUGE_f is garbled; it appears as 2 × (ROUGE_p + ROUGE_p)/(ROUGE_p + ROUGE_p), which is identically 2. It should be the standard harmonic mean, e.g., F = 2 · P · R / (P + R).
  2. [Section 3.4.2] The text says that evaluating 200 descriptions with nine metrics yields 1,600 computations, but 200 × 9 = 1,800; the same count appears to be repeated for the baseline, so the numbers should be corrected.
  3. [Eqs. (3) and (4)] The summation notation appears to run from i=0 to i=200, which would be 201 terms; with 200 test samples the index should run from 1 to 200 (or 0 to 199).
  4. [Table 7, Section 4.2] The text states that cosine similarity values from Eqs. (3)–(5) are 'not particularly informative,' yet Table 7 reports an average improvement of 22.27% for cosine similarity; the conditions under which Eq. (5) is applied to a metric with range [-1,1] should be clarified.
  5. [Section 5] The conclusion says 'six of these metrics demonstrate that ErgoChat achieves superior IC results for over 90% of the data,' but Table 7 lists seven metrics with improvement rates above 90% (ROUGE_r, ROUGE_f, BLEU, NIST, cosine similarity, METEOR, and SPICE).
  6. [Figure 3 caption, Section 3.3] The caption contains a grammatical error: 'A image from the fine-tuning partition' should be 'An image from the fine-tuning partition.'
  7. [Introduction, reproducibility note] The paper states the software and dataset 'will be publicly accessible' at the GitHub link; since reproducibility is listed as part of the contribution, the repository should be made available at the time of publication (or the current availability status should be stated).

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the paper reports a standard supervised fine-tuning evaluation with an appropriate same-architecture control; possible train/test overlap is a data-quality concern, not a circularity.

full rationale

This paper is an empirical machine-learning system paper, not a mathematical derivation, so the circularity patterns defined here (self-definitional equations, fitted parameters renamed as predictions, load-bearing self-citation chains, imported uniqueness theorems, ansatz smuggling, or renaming known results) do not apply to its central claims. The headline claim is that a VLM fine-tuned on the authors' ergonomic image-text dataset outperforms the same architecture without that fine-tuning, measured by VQA accuracy, nine IC metrics, and human expert preference. The comparison against the same architecture without fine-tuning on the same test partition is a standard and appropriate control, and the reported gains are empirical measurements rather than consequences of the definitions. The training and test sets are both drawn from free-to-use 'construction works' images and annotated under the same REBA-based protocol, which raises a legitimate generalization and possible-leakage concern if near-duplicates exist between the 1,700 fine-tuning and 200 testing images; however, this is a data-dependence / correctness risk and not a circularity, because the test labels are not themselves used to construct the fine-tuning objective as an algebraic identity. The self-citations to the authors' prior work (e.g., [13], [15], [20]) appear in related-work or methods context and are not load-bearing for the validity of the ErgoChat evaluation. No equation in the paper reduces to an input by construction, and no fitted parameter is renamed as a prediction. Accordingly, no specific circular step can be quoted, and the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claims rest on the validity of REBA-based labels, the representativeness of the online image source, and the expertise of the human evaluators. No new physical entities or hand-fitted constants are introduced; the model parameters are learned by standard fine-tuning.

assumptions (3)
  • domain assumption REBA risk scoring is a valid and sufficient ground truth for postural ergonomic risk in construction images.
    Dataset text labels are derived from REBA thresholds (medium or above considered risk), Section 3.2.
  • domain assumption Images collected by searching 'construction works' in free-to-use online libraries are representative of real construction worker postures.
    Both training and testing partitions are drawn from this single source, Section 3.2.
  • domain assumption Self-reported ergonomic knowledge of questionnaire respondents is sufficient to judge caption accuracy.
    Human evaluation section treats participants as experts, while only 10% self-identify as experts, Section 4.2.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ErgoChat: a Visual Query System for the Ergonomic Risk Assessment of Construction Workers." pith.science (2026). https://pith.science/paper/OOJKYOU7

@misc{pith2026241219954,
  author       = {Pith},
  title        = {Pith review of: ErgoChat: a Visual Query System for the Ergonomic Risk Assessment of Construction Workers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OOJKYOU7}},
  note         = {Machine review of arXiv:2412.19954}
}
read the original abstract

In the construction sector, workers often endure prolonged periods of high-intensity physical work and prolonged use of tools, resulting in injuries and illnesses primarily linked to postural ergonomic risks, a longstanding predominant health concern. To mitigate these risks, researchers have applied various technological methods to identify the ergonomic risks that construction workers face. However, traditional ergonomic risk assessment (ERA) techniques do not offer interactive feedback. The rapidly developing vision-language models (VLMs), capable of generating textual descriptions or answering questions about ergonomic risks based on image inputs, have not yet received widespread attention. This research introduces an interactive visual query system tailored to assess the postural ergonomic risks of construction workers. The system's capabilities include visual question answering (VQA), which responds to visual queries regarding workers' exposure to postural ergonomic risks, and image captioning (IC), which generates textual descriptions of these risks from images. Additionally, this study proposes a dataset designed for training and testing such methodologies. Systematic testing indicates that the VQA functionality delivers an accuracy of 96.5%. Moreover, evaluations using nine metrics for IC and assessments from human experts indicate that the proposed approach surpasses the performance of a method using the same architecture trained solely on generic datasets. This study sets a new direction for future developments in interactive ERA using generative artificial intelligence (AI) technologies.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

85 extracted references · 44 canonical work pages

  1. [1]

    Bureau of Labor Statistics, n.d

    Injuries, Illnesses, and Fatalities, U.S. Bureau of Labor Statistics, n.d. https://www.bls.gov/iif/factsheets/msds.htm (accessed December 3, 2023)

  2. [2]

    https://awcbc.org/en/statistics/ (accessed July 9, 2023)

    National Work Injury/Disease Statistics Program (NWISP), Association of Workers ' Compensation Boards of Canada, n.d. https://awcbc.org/en/statistics/ (accessed July 9, 2023)

  3. [3]

    https://osha.europa.eu/en/publications/msds-facts-and-figures-overview-prevalence-costs- and-demographics-msds-europe (accessed April 24, 2024)

    European Agency for Safety and Health at Work, Work- related musculoskeletal disorders: prevalence, costs and demographics in the EU, n.d. https://osha.europa.eu/en/publications/msds-facts-and-figures-overview-prevalence-costs- and-demographics-msds-europe (accessed April 24, 2024)

  4. [4]

    David, Ergonomic methods for assessing exposure to risk factors for work- related musculoskeletal disorders, Occup

    G.C. David, Ergonomic methods for assessing exposure to risk factors for work- related musculoskeletal disorders, Occup. Med. 55 (2005) 190 –199. https://doi.org/10.1093/occmed/kqi082

  5. [5]

    MassirisFernández, J.Á

    M. MassirisFernández, J.Á. Fernández, J.M. Bajo, C.A. Delrieux, Ergonomic risk assessment based on computer vision and machine learning, Comput. Ind. Eng. 149 (2020) 106816. https://doi.org/10.1016/j.cie.2020.106816

  6. [6]

    Jeong, J

    S. Jeong, J. Kook, CREBAS: Computer -Based REBA Evaluation System for Wood Manufacturers Using MediaPipe, Appl. Sci. 13 (2023) 938. https://doi.org/10.3390/app13020938

  7. [7]

    Barberi, M

    E. Barberi, M. Chillemi, F. Cucinotta, D. Milardi, M. Raffaele, F. Salmeri, F. Sfravara, Posture Interactive Self Evaluation Algorithm Based on Computer Vision, in: S. Gerbino, A. Lanzotti, M. Martorelli, R. Mirálbes Buil, C. Rizzi, L. Roucoules (Eds.), Adv. Mech. Des. Eng. Manuf. IV , Springer International Publishing, Cham, 2023: pp. 1516– 1526. https:/...

  8. [8]

    Nayak, E

    G.K. Nayak, E. Kim, Development of a fully automated RULA assessment system based on computer vision, Int. J. Ind. Ergon. 86 (2021) 103218. https://doi.org/10.1016/j.ergon.2021.103218

Show all 85 references
  1. [9]

    J. Seo, K. Yin, S. Lee, Automated Postural Ergonomic Assessment Using a Computer Vision- Based Posture Classification, in: Constr. Res. Congr. 2016, American Society of Civil Engineers, San Juan, Puerto Rico, 2016: pp. 809 –818. https://doi.org/10.1061/9780784479827.082

  2. [10]

    Plantard, H.P.H

    P. Plantard, H.P.H. Shum, A.-S. Le Pierres, F. Multon, Validation of an ergonomic assessment method using Kinect data in real workplace conditions, Appl. Ergon. 65 (2017) 562– 569. https://doi.org/10.1016/j.apergo.2016.10.015

  3. [11]

    Vignais, F

    N. Vignais, F. Bernard, G. Touvenot, J.-C. Sagot, Physical risk factors identification based on body sensor network combined to videotaping, Appl. Ergon. 65 (2017) 410 –417. https://doi.org/10.1016/j.apergo.2017.05.003

  4. [12]

    Donisi, G

    L. Donisi, G. Cesarelli, N. Pisani, A.M. Ponsiglione, C. Ricciardi, E. Capodaglio, Wearable Sensors and Artificial Intelligence for Physical Ergonomics: A Systematic Review of Literature, Diagnostics 12 (2022) 3048. https://doi.org/10.3390/diagnostics12123048

  5. [13]

    C. Fan, Q. Mei, X. Li, 3D pose estimation dataset and deep learning -based ergonomic risk assessment in construction, Autom. Constr. 164 (2024) 105452. https://doi.org/10.1016/j.autcon.2024.105452

  6. [14]

    Yenduri, M

    G. Yenduri, M. Ramalingam, G.C. Selvi, Y . Supriya, G. Srivastava, P.K.R. Maddikunta, G.D. Raj, R.H. Jhaveri, B. Prabadevi, W. Wang, A.V . Vasilakos, T.R. Gadekallu, GPT (Generative Pre-Trained Transformer)— A Comprehensive Review on Enabling Technologi es, Potential Applicati...

  7. [15]

    Aliasgari, C

    R. Aliasgari, C. Fan, X. Li, A. Golabchi, F. Hamzeh, MOCAP and AI -Based Automated Physical Demand Analysis for Workplace Safety, J. Constr. Eng. Manag. 150 (2024) 04024060. https://doi.org/10.1061/JCEMD4.COENG-13811

  8. [16]

    Vijayakumar, J

    R. Vijayakumar, J. Choi, Emerging Trends of Ergonomic Risk Assessment in Construction Safety Management: A Scientometric Visualization Analysis, Int. J. Environ. Res. Public. Health 19 (2022) 16120. https://doi.org/10.3390/ijerph192316120

  9. [17]

    Prince, K.B

    S.A. Prince, K.B. Adamo, M. Hamel, J. Hardt, S. Connor Gorber, M. Tremblay, A comparison of direct versus self -report measures for assessing physical activity in adults: a systematic review, Int. J. Behav. Nutr. Phys. Act. 5 (2008) 56. https://doi.org/10.1186/1479-5868-5-56

  10. [18]

    Crawford, The Nordic Musculoskeletal Questionnaire, Occup

    J.O. Crawford, The Nordic Musculoskeletal Questionnaire, Occup. Med. 57 (2007) 300–301. https://doi.org/10.1093/occmed/kqm036

  11. [19]

    Dawson, E.J

    A.P. Dawson, E.J. Steele, P.W. Hodges, S. Stewart, Development and Test–Retest Reliability of an Extended Version of the Nordic Musculoskeletal Questionnaire (NMQ-E): A Screening Instrument for Musculoskeletal Pain, J. Pain 10 (2009) 517 –526. https://doi.org/10.1016/j.jpain.2...

  12. [20]

    X. Li, S. Han, M. Gul, M. Al-Hussein, Automated Ergonomic Risk Assessment Based on 3D Visualization, in: Taipei, Taiwan, 2017. https://doi.org/10.22260/ISARC2017/0052

  13. [21]

    Guo, L.Y

    S.Y . Guo, L.Y . Ding, H.B. Luo, X.Y . Jiang, A Big-Data-based platform of workers' behavior: Observations from the field, Accid. Anal. Prev. 93 (2016) 299 –309. https://doi.org/10.1016/j.aap.2015.09.024

  14. [22]

    Marklin, J.R

    R.W. Marklin, J.R. Wilzbacher, Four Assessment Tools of Ergonomics Interventions: Case Study at an Electric Utility's Warehouse System, Am. Ind. Hyg. Assoc. J. 60 (1999) 777–784. https://doi.org/10.1080/00028899908984501

  15. [23]

    Lowe, P.G

    B.D. Lowe, P.G. Dempsey, E.M. Jones, Ergonomics assessment methods used by ergonomics professionals, Appl. Ergon. 81 (2019) 102882. https://doi.org/10.1016/j.apergo.2019.102882

  16. [24]

    Hignett, L

    S. Hignett, L. McAtamney, Rapid Entire Body Assessment (REBA), Appl. Ergon. 31 (2000) 201–205. https://doi.org/10.1016/S0003-6870(99)00039-3

  17. [25]

    McAtamney, E

    L. McAtamney, E. Nigel Corlett, RULA: a survey method for the investigation of work - related upper limb disorders, Appl. Ergon. 24 (1993) 91 –99. https://doi.org/10.1016/0003- 6870(93)90080-S

  18. [26]

    Fox, M.-L

    R.R. Fox, M.-L. Lu, E. Occhipinti, M. Jaeger, Understanding outcome metrics of the revised NIOSH lifting equation, Appl. Ergon. 81 (2019) 102897. https://doi.org/10.1016/j.apergo.2019.102897

  19. [27]

    Matthews, S.N

    J.D. Matthews, S.N. MacKinnon, W.J. Albert, M. Holmes, A. Patterson, Effects of moving environments on the physical demands of heavy materials handling operators, Int. J. Ind. Ergon. 37 (2007) 43–50. https://doi.org/10.1016/j.ergon.2006.09.018

  20. [28]

    Yunus, M.H

    M.N.H. Yunus, M.H. Jaafar, A.S.A. Mohamed, N.Z. Azraai, Md.S. Hossain, Implementation of Kinetic and Kinematic Variables in Ergonomic Risk Assessment Using Motion Capture Simulation: A Review, Int. J. Environ. Res. Public. Health 18 (2021) 8342. https://doi.org/10.3390/ijerph18168342

  21. [29]

    Colim, C

    A. Colim, C. Faria, A.C. Braga, N. Sousa, L. Rocha, P. Carneiro, N. Costa, P. Arezes, Towards an Ergonomic Assessment Framework for Industrial Assembly Workstations—A Case Study, Appl. Sci. 10 (2020) 3048. https://doi.org/10.3390/app10093048

  22. [30]

    Onofrejova, M

    D. Onofrejova, M. Balazikova, J. Glatz, Z. Kotianova, K. Vaskovicova, Ergonomic Assessment of Physical Load in Slovak Industry Using Wearable Technologies, Appl. Sci. 12 (2022) 3607. https://doi.org/10.3390/app12073607

  23. [31]

    Atici, D

    H. Atici, D. Gonen, A. Oral, B. Kaya, Ergonomic Analysis of an Assembly Line Using the AnyBody Modeling System, in: 2017. https://doi.org/10.11159/icmie17.125

  24. [32]

    Mortensen, M

    J. Mortensen, M. Trkov, A. Merryweather, Improved ergonomic risk factor assessment using opensim and inertial measurement units, in: Proc. 2018 IEEEACM Int. Conf. Connect. Health Appl. Syst. Eng. Technol., ACM, Washington DC, 2018: pp. 27– 28. https://doi.org/10.1145/3278576.3278589

  25. [33]

    Sekkay, Prevention of Work- Related Musculoskeletal Disorders supported by Artificial Intelligence, in: 2023

    F. Sekkay, Prevention of Work- Related Musculoskeletal Disorders supported by Artificial Intelligence, in: 2023. https://doi.org/10.54941/ahfe1002945

  26. [34]

    J. Cai, X. Li, X. Liang, W. Wei, S. Li, Construction Worker Ergonomic Assessment via LSTM- Based Multi-Task Learning Framework, in: Constr. Res. Congr. 2022, American Society of Civil Engineers, Arlington, Virginia, 2022: pp. 215– 224. https://doi.org/10.1061/9780784483961.023

  27. [35]

    Antwi-Afari, Y

    M.F. Antwi-Afari, Y . Qarout, R. Herzallah, S. Anwer, W. Umer, Y . Zhang, P. Manu, Deep learning-based networks for automated recognition and classification of awkward working postures in construction using wearable insole sensor data, Autom. Constr. 136 (2022) 104181. https:/...

  28. [36]

    J. Zhao, E. Obonyo, Applying incremental Deep Neural Networks-based posture recognition model for ergonomics risk assessment in construction, Adv. Eng. Inform. 50 (2021) 101374. https://doi.org/10.1016/j.aei.2021.101374

  29. [37]

    Akanmu, J

    A.A. Akanmu, J. Olayiwola, O. Ogunseiju, D. McFeeters, Cyber -physical postural training system for construction workers, Autom. Constr. 117 (2020) 103272. https://doi.org/10.1016/j.autcon.2020.103272

  30. [38]

    Antwi-Afari, H

    M.F. Antwi-Afari, H. Li, W. Umer, Y . Yu, X. Xing, Construction Activity Recognition and Ergonomic Risk Assessment Using a Wearable Insole Pressure System, J. Constr. Eng. Manag. 146 (2020) 04020077. https://doi.org/10.1061/(ASCE)CO.1943-7862.0001849

  31. [39]

    W. Umer, H. Li, Y . Yantao, M.F. Antwi-Afari, S. Anwer, X. Luo, Physical exertion modeling for construction tasks using combined cardiorespiratory and thermoregulatory measures, Autom. Constr. 112 (2020) 103079. https://doi.org/10.1016/j.autcon.2020.103079

  32. [40]

    J. Zhao, E. Obonyo, Convolutional long short -term memory model for recognizing construction workers' postures from wearable inertial measurement units, Adv. Eng. Inform. 46 (2020) 101177. https://doi.org/10.1016/j.aei.2020.101177

  33. [41]

    Zhang, M.M

    L. Zhang, M.M. Diraneyya, J. Ryu, C. Haas, E. Abdel- Rahman, Automated Monitoring of Physical Fatigue Using Jerk, in: Banff, AB, Canada, 2019. https://doi.org/10.22260/ISARC2019/0132

  34. [42]

    Y . Yu, H. Li, X. Yang, W. Umer, Estimating Construction Workers ' Physical Workload by Fusing Computer Vision and Smart Insole Technologies, in: Taipei, Taiwan, 2018. https://doi.org/10.22260/ISARC2018/0168

  35. [43]

    Nath, A.H

    N.D. Nath, A.H. Behzadan, Construction Productivity and Ergonomic Assessment Using Mobile Sensors and Machine Learning, in: Comput. Civ. Eng. 2017, American Society of Civil Engineers, Seattle, Washington, 2017: pp. 434– 441. https://doi.org/10.1061/9780784480847.054

  36. [44]

    Damilos, S

    S. Damilos, S. Saliakas, D. Karasavvas, E.P. Koumoulos, An Overview of Tools and Challenges for Safety Evaluation and Exposure Assessment in Industry 4.0, Appl. Sci. 14 (2024) 4207. https://doi.org/10.3390/app14104207

  37. [45]

    N. Rane, S. Choudhary, J. Rane, Integrating ChatGPT, Bard, and leading -edge generative artificial intelligence in building and construction industry: applications, framework, challenges, and future scope, SSRN Electron. J. (2023). https://doi.org/10.2139/ssrn.4645597

  38. [46]

    G. Yong, M. Liu, S. Lee, Automated Captioning for Ergonomic Problem and Solution Identification in Construction Using a Vision-Language Model and Caption Augmentation, in: Constr. Res. Congr. 2024, American Society of Civil Engineers, Des Moines, Iowa, 2024: pp. 709–718. https...

  39. [47]

    G. Yong, M. Liu, S. Lee, Explainable Image Captioning to Identify Ergonomic Problems and Solutions for Construction Workers, J. Comput. Civ. Eng. 38 (2024) 04024022. https://doi.org/10.1061/JCCEE5.CPENG-5744

  40. [48]

    Smetana, L

    M. Smetana, L. Salles De Salles, I. Sukharev, L. Khazanovich, Highway Construction Safety Analysis Using Large Language Models, Appl. Sci. 14 (2024) 1352. https://doi.org/10.3390/app14041352

  41. [49]

    Uddin, A

    S.M.J. Uddin, A. Albert, A. Ovid, A. Alsharef, Leveraging ChatGPT to Aid Construction Hazard Recognition and Support Safety Education and Training, Sustainability 15 (2023)

  42. [50]

    B. Yoo, J. Kim, S. Park, C.R. Ahn, T. Oh, Harnessing Generative Pre -Trained Transformers for Construction Accident Prediction with Saliency Visualization, Appl. Sci. 14 (2024) 664. https://doi.org/10.3390/app14020664

  43. [51]

    H. Chen, L. Hou, S. Wu, G. Zhang, Y . Zou, S. Moon, M. Bhuiyan, Augmented reality, deep learning and vision-language query system for construction worker safety, Autom. Constr. 157 (2024) 105158. https://doi.org/10.1016/j.autcon.2023.105158

  44. [52]

    J. Li, D. Li, C. Xiong, S. Hoi, BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation, (2022). https://doi.org/10.48550/ARXIV .2201.12086

  45. [53]

    J. Chen, D. Zhu, X. Shen, X. Li, Z. Liu, P. Zhang, R. Krishnamoorthi, V . Chandra, Y . Xiong, M. Elhoseiny, MiniGPT-v2: large language model as a unified interface for vision -language multi-task learning, (2023). https://doi.org/10.48550/ARXIV .2310.09478

  46. [54]

    Lecler, L

    A. Lecler, L. Duron, P. Soyer, Revolutionizing radiology with GPT -based models: Current applications, future possibilities and limitations of ChatGPT, Diagn. Interv. Imaging 104 (2023) 269–274. https://doi.org/10.1016/j.diii.2023.02.003

  47. [55]

    Papineni, S

    K. Papineni, S. Roukos, T. Ward, W. -J. Zhu, BLEU: a method for automatic evaluation of machine translation, in: Proc. 40th Annu. Meet. Assoc. Comput. Linguist. - ACL 02, Association for Computational Linguistics, Philadelphia, Pennsylvania, 2001: p. 3 11. https://doi.org/10.3...

  48. [56]

    Stefanini, M

    M. Stefanini, M. Cornia, L. Baraldi, S. Cascianelli, G. Fiameni, R. Cucchiara, From Show to Tell: A Survey on Deep Learning-Based Image Captioning, IEEE Trans. Pattern Anal. Mach. Intell. 45 (2023) 539–559. https://doi.org/10.1109/TPAMI.2022.3148210

  49. [57]

    Y . Fang, W. Wang, B. Xie, Q. Sun, L. Wu, X. Wang, T. Huang, X. Wang, Y . Cao, EV A: Exploring the Limits of Masked Visual Representation Learning at Scale, in: 2023 IEEECVF Conf. Comput. Vis. Pattern Recognit. CVPR, IEEE, Vancouver, BC, Canada, 2023: pp . 19358–19369. https:/...

  50. [58]

    Touvron, L

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C.C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V . Goswami, N. Goyal, A. Hartshorn, S. Hos...

  51. [59]

    Z. Peng, W. Wang, L. Dong, Y . Hao, S. Huang, S. Ma, F. Wei, Kosmos -2: Grounding Multimodal Large Language Models to the World, (2023). https://doi.org/10.48550/ARXIV .2306.14824

  52. [60]

    Schuhmann, R

    C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, A. Komatsuzaki, LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image- Text Pairs, (2021). https://doi.org/10.48550/ARXIV .2111.02114

  53. [61]

    Sharma, N

    P. Sharma, N. Ding, S. Goodman, R. Soricut, Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning, in: Proc. 56th Annu. Meet. Assoc. Comput. Linguist. V ol. 1 Long Pap., Association for Computational Linguistics, Melbourne, Australia...

  54. [62]

    J. Shawe -Taylor, Neural Information Processing Systems Foundation, eds., Advances in neural information processing systems 24: 25th Annual Conference on Neural Information Processing Systems 2011 ; December 12 - 15, 2011, Granada, Spain, Curran, Red Hook, NY , 2012

  55. [63]

    T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, C.L. Zitnick, Microsoft COCO: Common Objects in Context, in: D. Fleet, T. Pajdla, B. Schiele, T. Tuytelaars (Eds.), Comput. Vis. – ECCV 2014, Springer International Publishing, Cham, 2014: pp. 740–75...

  56. [64]

    Sidorov, R

    O. Sidorov, R. Hu, M. Rohrbach, A. Singh, TextCaps: A Dataset for Image Captioning with Reading Comprehension, in: A. Vedaldi, H. Bischof, T. Brox, J. -M. Frahm (Eds.), Comput. Vis. – ECCV 2020, Springer International Publishing, Cham, 2020: pp. 742– 758. https://doi.org/10.10...

  57. [65]

    Kazemzadeh, V

    S. Kazemzadeh, V . Ordonez, M. Matten, T. Berg, ReferItGame: Referring to Objects in Photographs of Natural Scenes, in: Proc. 2014 Conf. Empir. Methods Nat. Lang. Process. EMNLP, Association for Computational Linguistics, Doha, Qatar, 2014: pp. 787 –798. https://doi.org/10.311...

  58. [66]

    L. Yu, P. Poirson, S. Yang, A.C. Berg, T.L. Berg, Modeling Context in Referring Expressions, in: B. Leibe, J. Matas, N. Sebe, M. Welling (Eds.), Comput. Vis. – ECCV 2016, Springer International Publishing, Cham, 2016: pp. 69 –85. https://doi.org/10.1007/978-3-319-46475- 6_5

  59. [67]

    J. Mao, J. Huang, A. Toshev, O. Camburu, A. Yuille, K. Murphy, Generation and Comprehension of Unambiguous Object Descriptions, (2015). https://doi.org/10.48550/ARXIV .1511.02283

  60. [68]

    Krishna, Y

    R. Krishna, Y . Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y . Kalantidis, L.-J. Li, D.A. Shamma, M.S. Bernstein, L. Fei-Fei, Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations, Int. J. Comput. Vis. 12 3 (2017) 32– 73. https:...

  61. [69]

    Hudson, C.D

    D.A. Hudson, C.D. Manning, GQA: A New Dataset for Real- World Visual Reasoning and Compositional Question Answering, in: 2019 IEEECVF Conf. Comput. Vis. Pattern Recognit. CVPR, IEEE, Long Beach, CA, USA, 2019: pp. 6693 –6702. https://doi.org/10.1109/CVPR.2019.00686

  62. [70]

    Goyal, T

    Y . Goyal, T. Khot, D. Summers-Stay, D. Batra, D. Parikh, Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering, in: 2017 IEEE Conf. Comput. Vis. Pattern Recognit. CVPR, IEEE, Honolulu, HI, 2017: pp. 6325– 6334. https://doi.org/10.1...

  63. [71]

    Mishra, S

    A. Mishra, S. Shekhar, A.K. Singh, A. Chakraborty, OCR-VQA: Visual Question Answering by Reading Text in Images, in: 2019 Int. Conf. Doc. Anal. Recognit. ICDAR, IEEE, Sydney, Australia, 2019: pp. 947–952. https://doi.org/10.1109/ICDAR.2019.00156

  64. [72]

    Marino, M

    K. Marino, M. Rastegari, A. Farhadi, R. Mottaghi, OK -VQA: A Visual Question Answering Benchmark Requiring External Knowledge, in: 2019 IEEECVF Conf. Comput. Vis. Pattern Recognit. CVPR, IEEE, Long Beach, CA, USA, 2019: pp. 3190 –3199. https://doi.org/10.1109/CVPR.2019.00331

  65. [73]

    Schwenk, A

    D. Schwenk, A. Khandelwal, C. Clark, K. Marino, R. Mottaghi, A -OKVQA: A Benchmark for Visual Question Answering Using World Knowledge, in: S. Avidan, G. Brostow, M. Cissé, G.M. Farinella, T. Hassner (Eds.), Comput. Vis. – ECCV 2022, Springer Nature Switzerland, Cham, 2022: pp...

  66. [74]

    H. Liu, C. Li, Q. Wu, Y .J. Lee, Visual Instruction Tuning, (2023). https://doi.org/10.48550/ARXIV .2304.08485

  67. [75]

    Plummer, L

    B.A. Plummer, L. Wang, C.M. Cervantes, J.C. Caicedo, J. Hockenmaier, S. Lazebnik, Flickr30k Entities: Collecting Region -to-Phrase Correspondences for Richer Image- to- Sentence Models, in: 2015 IEEE Int. Conf. Comput. Vis. ICCV , IEEE, Santiago, Chile, 2015: pp. 2641–2649. ht...

  68. [76]

    Honovich, T

    O. Honovich, T. Scialom, O. Levy, T. Schick, Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor, in: Proc. 61st Annu. Meet. Assoc. Comput. Linguist. V ol. 1 Long Pap., Association for Computational Linguistics, Toronto, Canada, 2023: pp. 14409–14428. h...

  69. [77]

    https://unsplash.com/photos/a - man-in-a-blue-hoodie-working-on-a-fence-1fkvMVm7trU (accessed April 26, 2024)

    JSB Co., a man in a blue hoodie working on a fence, 2023. https://unsplash.com/photos/a - man-in-a-blue-hoodie-working-on-a-fence-1fkvMVm7trU (accessed April 26, 2024)

  70. [78]

    Jelinek, R.L

    F. Jelinek, R.L. Mercer, L.R. Bahl, J.K. Baker, Perplexity —a measure of the difficulty of speech recognition tasks, J. Acoust. Soc. Am. 62 (1977) S63– S63. https://doi.org/10.1121/1.2016299

  71. [79]

    Lin, ROUGE: A Package for Automatic Evaluation of Summaries, in: Text Summ

    C.-Y . Lin, ROUGE: A Package for Automatic Evaluation of Summaries, in: Text Summ. Branches Out, Association for Computational Linguistics, Barcelona, Spain, 2004. https://aclanthology.org/W04-1000 (accessed April 26, 2024)

  72. [80]

    Doddington, Automatic evaluation of machine translation quality using n- gram co - occurrence statistics, in: Proc

    G. Doddington, Automatic evaluation of machine translation quality using n- gram co - occurrence statistics, in: Proc. Second Int. Conf. Hum. Lang. Technol. Res. -, Association for Computational Linguistics, San Diego, California, 2002: p. 138. https://doi.org/10.3115/1289189.1289273

  73. [81]

    S. Pal, M. Chang, M.F. Iriarte, Summary Generation Using Natural Language Processing Techniques and Cosine Similarity, in: A. Abraham, N. Gandhi, T. Hanne, T. -P. Hong, T. Nogueira Rios, W. Ding (Eds.), Intell. Syst. Des. Appl., Springer International Publishing, Cham, 2022: p...

  74. [82]

    Salihu, I.P

    S.A. Salihu, I.P. Onyekwere, M.A. Mabayoje, H.A. Mojeed, Performance Evaluation of Manhattan and Euclidean Distance Measures For Clustering Based Automatic Text Summarization, FUOYE J. Eng. Technol. 4 (2019). https://doi.org/10.46792/fuoyejet.v4i1.316

  75. [83]

    Lavie, A

    A. Lavie, A. Agarwal, Meteor: an automatic metric for MT evaluation with high levels of correlation with human judgments, in: Proc. Second Workshop Stat. Mach. Transl. - StatMT 07, Association for Computational Linguistics, Prague, Czech Republic, 200 7: pp. 228–231. https://d...

  76. [84]

    Anderson, B

    P. Anderson, B. Fernando, M. Johnson, S. Gould, SPICE: Semantic Propositional Image Caption Evaluation, in: B. Leibe, J. Matas, N. Sebe, M. Welling (Eds.), Comput. Vis. – ECCV 2016, Springer International Publishing, Cham, 2016: pp. 382– 398. https://doi.org/10.1007/978-3-319-...

  77. [7121]

    https://doi.org/10.3390/su15097121

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.