REVIEW 3 major objections 6 minor 21 references
Supporting Construction Worker Well-Being with a Multi-Agent Conversational AI System
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper claims that a multi-agent conversational AI system with distinct roles and retrieval-augmented domain knowledge outperforms a generic single-agent chatbot for construction-worker well-being support, reporting 18% higher…
desk verdict A promising pilot for a real need, but the experimental design cannot attribute the gains to multi-agent orchestration—read it as a system-level comparison, not a mechanism test. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the multi-agent orchestration layer combined with persona-specific retrieval-augmented generation (RAG). Each agent has its own prompt, personality, conversation goals, and a private vector-database knowledge source, while the conversation history is shared across agents. An orchestration module decides which agents should respond, and can run them sequentially or in parallel to simulate natural group conversation. This design lets agents support, complement, or disagree with one another, which the authors credit for producing more practical, multi-perspective advice and a stronger sense of human-like presence than the single-agent baseline.
What would settle it
Run the same within-subjects comparison with a sample of actual construction workers (or a much larger, more diverse participant group) and measure the same SUS, self-determination, social presence, and trust scales; if the multi-agent system no longer significantly outperforms the single-agent baseline, the central claim would be refuted. A behavioral follow-up could track whether workers choose to use the multi-agent system repeatedly or act on its advice.
Extended reading notes
Core claim
The central claim is that adding multiple specialized, persona-carrying AI agents with retrieval-augmented domain knowledge improves the user experience of a conversational support tool for construction-worker scenarios compared with a single generic LLM chatbot. In the reported study, the proposed system outperformed the baseline by 18% in usability, 40% in self-determination, and 60% in both social presence and trust, with the most striking gain in relatedness (from a mean of 2.08 to 3.83). The authors attribute these results to the agents' distinct identities, professional profiles, and domain-specific expertise, which made the interaction feel more like consulting several real experts and less like a generic text generator.
Load-bearing premise
The load-bearing premise is that 12 participants role-playing assigned construction-worker scenarios give evaluations that generalize to how real construction workers would experience the system.
Editorial extensions
If this is right
- If the reported effect sizes hold, a multi-agent configuration could make LLM-based worker-support tools substantially more usable and trustworthy than generic chatbots.
- The strongest gains in relatedness suggest that multiple personas address social connection and loneliness more than the content quality of a single chatbot alone.
- The configurable agent interface could let non-experts tailor AI support for other industries by changing agent roles and uploading relevant documents.
- The system's group-conversation format may reduce the perceived stigma of seeking help, since users are talking with a supportive team rather than a single authority figure.
- The approach provides a template for evaluating conversational AI in worker well-being through standardized usability and psychological-needs scales.
Reading between the lines
- The comparison conflates two variables: the multi-agent architecture and the added domain knowledge from RAG; a future study could give the single-agent baseline the same OSHA and labor documents to separate the effect of personas from the effect of knowledge.
- Because the participants were role-playing, the strong trust and social-presence results might not transfer to real construction workers, who face different stigma, language, and safety pressures; a field test with actual workers would be the natural next step.
- The same multi-persona design could plausibly improve perceived trust and relatedness in other high-stigma help-seeking contexts, such as counseling, addiction support, or financial advice, if the personas are chosen to match the user's community.
- A behavioral falsification test would measure whether workers actually act on the system's advice or return to it repeatedly, rather than relying only on self-reported trust and adoption intent.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a conversational multi-agent AI system for construction worker well-being, combining a multi-agent orchestration mechanism, retrieval-augmented generation (RAG) from external documents, and distinct agent personas (OSH Specialist, HR Advisor, Worker Peer). The system is evaluated in a within-subjects user study with 12 participants who role-played construction-worker scenarios, comparing the proposed system against a single-agent GPT-4o chatbot baseline. The authors report significantly higher SUS scores, basic psychological need satisfaction (autonomy, competence, relatedness), social presence, and trust for the proposed system, and conclude that the multi-agent design is responsible for these improvements.
Significance. If the findings are valid, the paper addresses a genuinely important and underexplored application area: conversational AI for mental-health and safety support in construction. The system implementation is concrete and described in adequate detail, and the authors make their user-study data available in figures and tables. The claim that a multi-agent LLM system with domain knowledge and personas improves usability, psychological need satisfaction, social presence, and trust relative to a generic chatbot is plausible and practically relevant. However, the experimental design confounds several independent variables, so the results cannot currently support the central attribution to the multi-agent architecture. The work is best read as a promising system demonstration with a preliminary user study, rather than a rigorous causal test of the multi-agent component.
major comments (3)
- [§3.5.1 and §3.4] The abstract's percentage improvements are not consistently derived from the reported means. The SUS improvement of 18% matches (84.58 vs 71.88). For self-determination, the reported subscale means give relative increases of 18% (autonomy: 4.14 vs 3.50), 19% (competence: 3.86 vs 3.25), and 84% (relatedness: 3.83 vs 2.08); the overall mean of the three subscales increases by 34%, not 40%. The 40% figure appears to be the unweighted average of the three subscale percentage increases. For social presence and trust, the 60% figure matches the average of the three item means (3.78 vs 2.36), but the trust item itself (Q2, 'the answer is practical') increases only by 37% (4.00 vs 2.92). The manuscript should report which composite or item is used for each headline percentage, and should provide effect sizes and confidence intervals to avoid overstating the magnitude of the effects.
- [§4.2 and §4.3] The statistical analysis does not account for multiple comparisons across the many outcome measures (SUS, three self-determination subscales, three social-presence/trust items) with only 12 participants. For example, the autonomy comparison yields p = 0.066, which is not significant at the conventional 0.05 level, yet the text describes it as 'marginally significant' and groups it with the significant results. No correction for multiple testing is applied, and no effect sizes or confidence intervals are reported. Given the small sample and the number of tests, the claim that the system 'significantly outperforms' on all dimensions should be tempered. At minimum, the authors should report exact p-values for all tests (they do for some but not all), apply a simple correction such as Holm-Bonferroni, and report effect sizes such as Cohen's dz for the within-subjects comparisons.
- [§3.5.3 and §5] The external validity of the evaluation is limited by the participant sample and the role-play design. The study recruited 12 participants who are not described as construction workers; they role-played predefined scenarios (Section 3.5.2) rather than drawing on their own lived experiences. The authors acknowledge this only in the conclusion ('small-sized participant group and predefined scenarios'), but the abstract and introduction frame the contribution as supporting construction workers. The role-play nature is particularly consequential for measures like relatedness and trust, which may be influenced by participants' identification with the scenario and by their prior experience with chatbots. The limitations section should be expanded and moved to a more prominent position, and the concrete claims about supporting construction workers should be restricted to 'a preliminary evaluation with role-played scenarios' until a study with actual workers is conducted.
minor comments (6)
- [Abstract] The phrase '60% in social presence and trust' is ambiguous; the results section reports three separate items and the percentage gain for the trust item is 37%, not 60%. Please clarify which composite is being quoted.
- [§3.4] The orchestration system 'randomly groups agents' and uses randomized ordering; this introduces variability in the user experience across trials. The paper should justify this design choice and discuss how the randomness might affect the measured outcomes.
- [§4.3.1] The three-item social presence and trust scale is described as having 'excellent reliability' (α = 0.911) and 'good reliability' (α = 0.817) for the two conditions, but with n = 12, Cronbach's alpha estimates are imprecise. Please report a confidence interval or note the small sample size in the interpretation.
- [References] Several references are incomplete or inconsistently formatted. For example, 'Lewis and Sauro 2018' is cited as 'Item benchmarks for the system.' 13 (3), which is missing the article title and journal name; 'Shu et al. 2024' is given as an arXiv preprint without a DOI. Please ensure all references follow a consistent and complete style.
- [Figures] Figures 4, 5, and 6 would benefit from explicit error bars on the bar charts, since the text reports means and standard deviations; currently the error bars are missing or not clearly labeled.
- [§4.1 and §4.2] Participant quotations are presented as support for the quantitative findings, but the open-ended responses are not analyzed systematically (e.g., no thematic coding). This is acceptable as illustration, but the text should avoid implying that the quotations constitute a formal qualitative analysis.
Circularity Check
No circularity: the reported gains are empirical outcomes of a randomized within-subjects user study, not consequences of how the system, scale, or baseline are defined.
full rationale
This paper makes no formal derivation claim. Its central result is a comparative user study in which 12 participants interacted with a proposed multi-agent RAG system and with a single-agent GPT-4o baseline, and the paper reports the measured differences in SUS, self-determination, social presence, and trust. No parameter is fitted to the outcome data, no prediction is constructed from the same measurements it claims to predict, and no uniqueness theorem or load-bearing self-citation is invoked. The authors use standard instruments (SUS and the La Guardia basic psychological needs scale) plus a three-item custom scale, but the custom scale is not calibrated to the observed results and therefore does not force the reported differences. The baseline differs from the treatment in multiple dimensions (multi-agent orchestration, predefined personas, and RAG-based external knowledge), which is a genuine internal-validity concern about attributing causality to any single component, but it is not circularity under the stated criteria: the treatment is not defined in terms of the outcome, and the improvement is not an algebraic identity. The acknowledged small sample and role-play scenarios limit generalizability but do not make the derivation circular. Accordingly, the paper warrants a score of 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The custom 3-item Perceived Social Presence and Trust scale (Q1-Q3) is a valid measure of social presence, trust, and adoption intent.
- domain assumption Role-play responses by 12 non-worker participants approximate how real construction workers would react to the system.
- ad hoc to paper The three selected agent personas (OSH Specialist, HR Advisor, Worker Peer) cover the most relevant support needs of construction workers.
- domain assumption GPT-4o behavior is stable across conditions and orderings, and the randomized exposure order fully controls for learning effects.
Cite this review
Pith. "Pith review of Supporting Construction Worker Well-Being with a Multi-Agent Conversational AI System." pith.science (2026). https://pith.science/paper/FVO7CGUS
@misc{pith2026250607997,
author = {Pith},
title = {Pith review of: Supporting Construction Worker Well-Being with a Multi-Agent Conversational AI System},
year = {2026},
howpublished = {\url{https://pith.science/paper/FVO7CGUS}},
note = {Machine review of arXiv:2506.07997}
}
read the original abstract
The construction industry is characterized by both high physical and psychological risks, yet supports of mental health remain limited. While advancements in artificial intelligence (AI), particularly large language models (LLMs), offer promising solutions, their potential in construction remains largely underexplored. To bridge this gap, we developed a conversational multi-agent system that addresses industry-specific challenges through an AI-driven approach integrated with domain knowledge. In parallel, it fulfills construction workers' basic psychological needs by enabling interactions with multiple agents, each has a distinct persona. This approach ensures that workers receive both practical problem-solving support and social engagement, ultimately contributing to their overall well-being. We evaluate its usability and effectiveness through a within-subjects user study with 12 participants. The results show that our system significantly outperforms the single-agent baseline, achieving improvements of 18% in usability, 40% in self-determination, 60% in social presence, and 60% in trust. These findings highlight the promise of LLM-driven AI systems in providing domain-specific support for construction workers.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Email: yang2352@purdue.edu 2Department of Computer Science, Purdue Univ., West Lafayette, IN, USA
Joint CSCE Construction Specialty & CRC Conference 2025 Conférence conjointe spécialisée en construction de la SCGC et CRC-2025 Montreal, Quebec July 28-31, 2025 / 28-31 juillet 2025 CON-172-1 Supporting Construction Worker Well-Being with a Multi-Agent Conversational AI System Fan Yang1; Yuan Tian2; and Jiansong Zhang, PhD., A.M.ASCE3 1School of Construc...
work page 2025
-
[2]
INTRODUCTION The construction industry is characterized by a high incidence of worker injuries and a prevalence of mental health challenges (Langdon and Sawang 2018; Ross et al. 2022). Reports indicate that the suicide rate among construction workers is 44% higher than the national average, underscoring the severity of the problem (Campbell and Gunning 20...
work page 2018
-
[3]
BACKGROUND 2.1 Emotional Health Support for Construction Workers The construction industry is characterized by high physical risks, demanding workloads, and job insecurity, all of which contribute to significant mental health challenges among workers. Studies indicate that construction workers experience elevated rates of stress, anxiety, depression, and ...
work page 2004
-
[4]
METHODOLOGY 3.1 Overview We experimented on system deployment, scenario-based interactions, and a user study to evaluate our proposed system. Figure 1 illustrates the proposed RAG-based conversational multi-agent system. User messages are processed by multiple collaborative agents that leverage a vector database and configurable external documentation. Ea...
work page 1996
-
[6]
The three items aim to capture users perceived social presence (i.e., whether the AI feels human-like), trust (i.e., whether the answer is practical), in the system's responses, and intend to adopt the system for future use. To assess the internal consistency of this scale, we conducted Cronbach’s Alpha reliability analysis using data collected from the t...
work page 2011
-
[10]
Mental health stigma and wellbeing among commercial construction workers: A mixed methods study
“Mental health stigma and wellbeing among commercial construction workers: A mixed methods study.” Journal of Occupational and Environmental Medicine, 62 (8): e423. https://doi.org/10.1097/JOM.0000000000001929. Fitzpatrick, K. K., A. Darcy, and M. Vierhile
-
[13]
“The effectiveness of organisational-level workplace mental health interventions on mental health and wellbeing in construction workers: A systematic review and recommended research agenda.” PLOS ONE, 17 (11): e0277114. Public Library of Science. https://doi.org/10.1371/journal.pone.0277114. Gunning, J. g., and E. Cooke
-
[14]
The influence of occupational stress on construction professionals
“The influence of occupational stress on construction professionals.” Building Research & Information, 24 (4): 213–221. Routledge. https://doi.org/10.1080/09613219608727532. La Guardia, J. G., R. M. Ryan, C. E. Couchman, and E. L. Deci
Show all 21 references
-
[15]
Construction workers’ well-being: What leads to depression, anxiety, and stress?
“Construction workers’ well-being: What leads to depression, anxiety, and stress?” Journal of Construction Engineering and Management, 144 (2): 04017100. American Society of Civil Engineers. https://doi.org/10.1061/(ASCE)CO.1943-7862.0001406. Lewis, J. R., and J. Sauro
-
[18]
Suicidal ideation and related factors in construction industry apprentices
“Suicidal ideation and related factors in construction industry apprentices.” Journal of Affective Disorders, 297: 294–300. https://doi.org/10.1016/j.jad.2021.10.073. Saka, A. B., L. O. Oyedele, L. A. Akanbi, S. A. Ganiyu, D. W. M. Chan, and S. A. Bello
2021 doi
-
[19]
Conversational artificial intelligence in the AEC industry: A review of present status, challenges and opportunities
“Conversational artificial intelligence in the AEC industry: A review of present status, challenges and opportunities.” Advanced Engineering Informatics, 55: 101869. https://doi.org/10.1016/j.aei.2022.101869. Shu, R., N. Das, M. Yuan, M. Sunkara, and Y. Zhang
2022
-
[1996]
I don’t think the baseline gave me any specific suggestions that I can take directly. The response from the baseline is too general, and I don’t feel any humanity
we normalize the raw SUS score to a 100-point scale. Figure 4 presents a comparison of the normalized SUS scores between the baseline system and our proposed system. The statistical analysis reveals that users rated our system (Mean = 84.58, SD = 8.95) significantly higher tha...
2018
-
[2011]
Making sense of Cronbach’s alpha
“Making sense of Cronbach’s alpha.” Int J Med Educ, 2: 53–55. https://doi.org/10.5116/ijme.4dfb.8dfd
-
[2012]
Constructing masculinity in the building trades: ‘most jobs in the construction industry can be done by women.’
“Constructing masculinity in the building trades: ‘most jobs in the construction industry can be done by women.’” Gender, Work & Organization, 19 (6): 654–676. https://doi.org/10.1111/j.1468-0432.2010.00551.x. Nwaogu, J. M., A. P. C. Chan, C. K. H. Hon, and A. Darko
2010
-
[2017]
Delivering cognitive behavior therapy to young adults with symptoms of depression and anxiety using a fully automated conversational agent (woebot): A randomized controlled trial
“Delivering cognitive behavior therapy to young adults with symptoms of depression and anxiety using a fully automated conversational agent (woebot): A randomized controlled trial.” JMIR Mental Health, 4 (2): e7785. https://doi.org/10.2196/mental.7785. Gómez-Salgado, C., J. C....
-
[2018]
Multi-agent systems: A survey
“Multi-agent systems: A survey.” IEEE Access, 6: 28573–28593. https://doi.org/10.1109/ACCESS.2018.2831228. CON-172-10 Eyllon, M., S. P. Vallas, J. T. Dennerlein, S. Garverich, D. Weinstein, K. Owens, and A. K. Lincoln
2018
-
[2019]
Review of global mental health research in the construction industry: A science mapping approach
“Review of global mental health research in the construction industry: A science mapping approach.” Engineering, Construction and Architectural Management, 27 (2): 385–410. Emerald Publishing Limited. https://doi.org/10.1108/ECAM-02-2019-0114. Rizk, Y., M. Awad, and E. W. Tunstel
-
[2020]
Strategies to improve mental health and well-being within the UK construction industry
“Strategies to improve mental health and well-being within the UK construction industry.” Proceedings of the Institution of Civil Engineers - Management, Procurement and Law, 173 (2): 64–74. ICE Publishing. https://doi.org/10.1680/jmapl.19.00020. Chaaban, Y., and C. Müller-Schloer
-
[2022]
Consensus in multi-agent systems: A review
“Consensus in multi-agent systems: A review.” Artif. Intell. Rev., 55 (5): 3897–3935. https://doi.org/10.1007/s10462-021-10097-x. Boiko, D. A., R. MacKnight, and G. Gomes
-
[2023]
Stress, fear, and anxiety among construction workers: a systematic review
“Stress, fear, and anxiety among construction workers: a systematic review.” Front. Public Health, 11: 1226914. https://doi.org/10.3389/fpubh.2023.1226914. Greiner B. A., Leduc C., O’Brien C., Cresswell-Smith J., Rugulies R., Wahlbeck K., Abdulla K., Amann B. L., Pashoja A. C....
2023
-
[2024]
Towards effective GenAI multi-agent collaboration: Design and evaluation for enterprise applications
“Towards effective GenAI multi-agent collaboration: Design and evaluation for enterprise applications.” arXiv preprint arXiv:2412.05449. Tavakol, M., and R. Dennick
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.