REVIEW 3 major objections 6 minor 7 references
OSU-Wing PIC Phase I Evaluation: Baseline Workload and Situation Awareness Results
T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Raising the number of supervised delivery UAS from about 30 to more than 75 did not raise pilot workload in simulator trials.
desk verdict Good baseline data from a real system; the abstract's 'debunk' claim outruns the design. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The evaluation's central object is a logarithmic workload model, $w = a(1 + r \ln n)$, which replaces the linear per-UAS workload assumption used in conventional modeling tools. Here $n$ is the number of active UAS, $a$ is the workload for a single UAS, and $r$ is a fitted logarithmic rate set to 0.5 for the delivery use case. This model carries the argument by predicting diminishing workload growth as UAS count increases, and the empirical result—flat physiological workload across UAS counts—is interpreted as confirming that sublinear relationship.
What would settle it
Rerun the comparison with condition order counterbalanced and at least 10–20 pilots per condition; if mean physiological workload rises reliably above the low-UAS baseline once active UAS exceed roughly 75, the central claim fails.
Extended reading notes
Core claim
The paper's central claim is that a single pilot in command can supervise substantially more autonomous delivery UAS without measurable performance cost, at least within the tested range. Increasing the number of active UAS by roughly 2.5 times left physiological workload estimates essentially flat, with overall means near 29 on a 0–100 scale and never approaching the overload threshold of 60; situation-awareness accuracy stayed at or above 92 percent. The authors argue this contradicts the fan-out theory developed for ground robots in the early 2010s, and they conclude that other factors—such as area size, UAS density, and geographical diversity—deserve more attention than raw UAS count.
Load-bearing premise
The evaluation assumes that six professional pilots completing two 60-minute sessions, with the high-UAS condition always second, provide enough power and comparability to detect a real workload effect of UAS count.
Editorial extensions
If this is right
- Operators could treat the tested range (roughly 25–30 to 75–90 active UAS) as one where pilot workload does not automatically cap fleet size.
- The results provide a baseline for the planned follow-up phase, which shifts attention to area size, UAS density, and geographical diversity as the factors likely to matter.
- Interface and training emphasis should remain on monitoring and exception response rather than per-UAS control, since the ADS-B display absorbed most visual attention and interactions.
- Workload modeling for highly autonomous multi-UAS operations should use sublinear growth in workload with UAS count rather than a linear per-UAS cost.
Reading between the lines
- If the flat-workload result reproduces, the binding constraint on how many delivery drones one pilot can supervise is likely airspace or interface complexity rather than raw UAS count.
- The logarithmic model predicts a gradual eventual rise in workload as UAS count keeps increasing, so the cleanest extension is to probe much larger fleet sizes and locate that crossover.
- Because every pilot flew the high-UAS condition second, a counterbalanced or between-subjects replication is the direct way to separate learning effects from any real fan-out cost.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a Phase I human-factors evaluation with six Wing-trained PICs in a high-fidelity simulated delivery-UAS environment. Three conditions were compared (10:Low, 24:Low, and 24:High), manipulating the number of nests and the number of active UAS, with nominal, DAA (single and double crewed aircraft), and adverse-weather tasks. The study collected objective physiological workload estimates from OSU's multi-dimensional workload algorithm, eye-tracking fixations, display-interaction logs, subjective in-situ workload, and SA probes. The results show overall workload in the normal-to-underload range, SA above 92% correct across trials, and no statistically reliable overall workload differences among conditions. The abstract and executive summary interpret these findings as debunking the theory that increasing the number of UAS degrades pilot performance.
Significance. If treated as a descriptive baseline, the dataset is valuable: it uses experienced, Part 107-certified PICs, Wing's actual delivery interface in a high-fidelity simulator, and a rich set of objective measures (physiological workload, eye tracking, interaction coding, SA probes). The reporting of descriptive statistics and the detailed AOI and interaction analyses are careful and will inform Phase II and regulatory conversations. However, the strong causal null claim ('debunk') is not supported by the design. The 24:High condition was always run second, the X:Low cells have n=3, the active-UAS count varied by week, and the large-df ANOVAs are pseudo-replicated. The paper's contribution is therefore best evaluated as a baseline study with descriptive findings, not as a test of the fanout theory.
major comments (3)
- [2.1, Figure 2; Executive Summary; Section 1] The central claim that increasing the number of active UAS is not detrimental is confounded by session order. All six pilots completed 24:High in the second session, after completing either 10:Low or 24:Low in the first session (Section 2.1, Figure 2). The paper itself attributes the lower subjective workload and improved SA in 24:High to learning effects (Sections 3.3.4 and 3.5), and Table 12 shows first-session workload declining across Nominal #1 to Nominal #3 within the X:Low trials. A non-significant difference between the first-session X:Low and second-session 24:High conditions cannot be attributed to the number of UAS rather than to practice and familiarization. The wording 'debunk' is also at odds with the paper's stated hypothesis that these factors were predicted to have little impact (Section 1). This claim should be removed or substantially qualified, with the study presented as an inconclusive baseline.
- [3.3.3.2, 3.3.3.3; 2.4] The ANOVAs for the active-UAS and nest comparisons are reported with very high degrees of freedom (e.g., F(531, 2124) and F(531, 2655) in Sections 3.3.3.2 and 3.3.3.3), implying that time points within each 10-minute task were treated as repeated measures. With only three pilots per X:Low condition and six in 24:High, the participant-level error term is not properly represented; these tests are pseudo-replicated and cannot support generalizable null conclusions. The a-priori power analysis in Section 2.4 (n=6, f=0.25) is not applicable to the between-subject comparison of 10:Low versus 24:Low, each with n=3. The non-significant results in Section 3.3.2 should be reported as descriptive evidence only, and the authors should either fit mixed models with participant as a random effect or explicitly refrain from inferential claims about the number of UAS.
- [3.1, Table 6; Abstract] The 24:High manipulation was not internally consistent. The number of active UAS differed substantially by week (mean 90.33 in week 1 vs. 62.33 in week 2, Table 6), and the authors state they could not identify the cause. This means the 24:High condition was not a controlled 'high' UAS condition, and the comparison with X:Low aggregates across two different active-UAS levels, as well as different weather and DAA scripts. The abstract's claim that the results 'debunk' the theory is therefore unsupported even if the statistical issues in the previous comment were resolved. At minimum, the paper should restrict conclusions to the specific UAS ranges observed and acknowledge that the week-to-week variability limits the manipulation.
minor comments (6)
- [Section 3] The text contains a typo: 'ANOV A' should be 'ANOVA'. A global pass for such mechanical errors is needed.
- [Section 2.2] There is a grammatical error: 'and the a Shure microphone is worn as a headset' should read 'and a Shure microphone is worn as a headset.'
- [Section 3.3.4] The sentence 'There was virtually no difference between and withing the remaining component results' contains a typo ('withing' should be 'within').
- [Section 3.2.3.1] The sentence 'A 322% increase in fixation durations on AOI 2,2 existed as compared to the Nominal #1 task, with a 202% increase for the adjacent 2,2 AOI' is confusing because the adjacent AOI is also identified as 2,2; please clarify which AOI was intended.
- [Table 15 (TOC)] The table of contents lists two entries for Table 15 with slightly different titles ('The normalized in situ overall workload results descriptive statistics by trial' and 'The normalized in situ overall workload descriptive statistics by trial'), but only one Table 15 appears in the text. Please harmonize the numbering and titles.
- [Section 2.5.1, Eq. (3)] The derivation of the logarithmic workload model is clear, but the manuscript should more explicitly state that the fitted rate r=0.5 is based on prior work and is not validated against the Phase I data; this would avoid any impression that the model is used as empirical evidence.
Circularity Check
No significant circularity: the workload and situation-awareness results are empirical outcomes, and the IMPRINT model is not used as evidence for the main conclusion.
full rationale
The paper's central claim that increasing the number of UAS and nests did not degrade pilot performance rests on measured physiological workload estimates, subjective workload ratings, interaction counts, eye-tracking data, and SA probe accuracy. These are data collected from the evaluation, not quantities defined in terms of the conclusion. The only potentially self-referential input is the IMPRINT Pro workload model, whose logarithmic rate r=0.5 is set from the authors' prior OFFSET and ASSURE work (Section 2.5.1), but the manuscript explicitly does not use that model's predictions as evidence for the empirical non-effect; Section 4 states that the model's predicted workload was substantially higher than the measured results and attributes this to modeling limitations. The OSU workload estimation algorithm is a measurement method rather than a fitted predictor of the condition effect, and no equation in the paper reduces the empirical outcome to an input. The SA scoring may be lenient in nominal tasks, but that is a construct-validity concern rather than a circular derivation. The confounded session order and small per-condition sample are experimental-validity risks, not circularity. No circular step was found, so the score is 0.
Assumptions & free parameters
free parameters (2)
- IMPRINT Pro logarithmic rate r =
0.5
- PIC ratio simulator parameters =
45 (10:Low), 35 (24:Low), 200 (24:High)
assumptions (4)
- domain assumption Workload can be modeled as w = a(1 + r ln n) for multi-UAS supervision.
- ad hoc to paper The physiological sensor streams, when processed through OSU's multi-dimensional workload estimation algorithm, produce a valid scale of workload with the stated underload/normal/overload thresholds.
- domain assumption The 60-minute simulated trials with scripted encounters and a static-slide weather display adequately represent real nominal operations and unexpected events for workload and SA measurement.
- domain assumption The six participants, randomly split into 10:Low and 24:Low for session 1, provide enough power to detect a meaningful fanout effect.
invented entities (1)
-
Adapted PIC Reference Sheet and Procedures Manual and scripted aircraft injection events
Cite this review
Pith. "Pith review of OSU-Wing PIC Phase I Evaluation: Baseline Workload and Situation Awareness Results." pith.science (2026). https://pith.science/paper/ZAXNOIR2
@misc{pith2026241118750,
author = {Pith},
title = {Pith review of: OSU-Wing PIC Phase I Evaluation: Baseline Workload and Situation Awareness Results},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZAXNOIR2}},
note = {Machine review of arXiv:2411.18750}
}
read the original abstract
The common theory is that human pilot's performance degrades when responsible for an increased number of uncrewed aircraft systems (UAS). This theory was developed in the early 2010's for ground robots and not highly autonomous UAS. It has been shown that increasing autonomy can mitigate some performance impacts associated with increasing the number of UAS. Overall, the Oregon State University-Wing collaboration seeks to understand what factors negatively impact a pilot's ability to maintain responsibility and control over an assigned set of active UAS. The Phase I evaluation establishes baseline data focused on the number of UAS and the number of nests increase. This evaluation focuses on nominal operations as well as crewed aircraft encounters and adverse weather changes. The results demonstrate that the pilots were actively engaged and had very good situation awareness. Manipulation of the conditions did not result in any significant differences in overall workload. The overall results debunk the theory that increasing the number of UAS is detrimental to pilot's performance.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
E. J. Bass, R. Amey, J. Glavan, T. Read, N. Raghunath, C. A. Sanchez, K. Silas, T. Haritos, and J. A. Adams, “Participant characteristics in human-in-the-loop studies with multiple unmanned vehicles including aircraft,” in Human Factors and Ergonomics Annual Meeting, 2022
work page 2022
-
[2]
A26: Establish pilot proficiency requirements multi-uas components, literature review,
E. Bass, J. A. Adams, C. A. Sanchez, and T. Haritos, “A26: Establish pilot proficiency requirements multi-uas components, literature review,” F AA’s ASSURE Center of Ex- cellence, Tech. Rep., February 2021
work page 2021
-
[3]
Measures of attention in autonomous and semi-autonomous multi-vehicle supervision,
J. J. Glavan, E. J. Bass, C. A. Sanchez, T. Read, and J. A. Adams, “Measures of attention in autonomous and semi-autonomous multi-vehicle supervision,” in IEEE International Conference on Human-Machine Systems, 2022
work page 2022
-
[4]
A26 A11L.UA V.74 : Establish pilot proficiency requirements multi- UAS components, final report,
J. A. Adams, P. Uriarte, C. A. Sanchez, T. Read, J. Glavan, E. J. Bass, T. Hari- tos, and K. Silas, “A26 A11L.UA V.74 : Establish pilot proficiency requirements multi- UAS components, final report,” F AA’s ASSURE Center of Excellence, Tech. Rep. A26 A11L.UA V.74, November 2022
work page 2022
-
[5]
The development of a short domain-general measure of working memory capacity,
F. L. Oswald, S. T. McAbee, T. S. Redick, and D. Z. Hambrick, “The development of a short domain-general measure of working memory capacity,” Behavior research methods, vol. 47, pp. 1343–1355, 2015
work page 2015
-
[6]
E. T. Service, J. W. French, R. B. Ekstrom, and L. A. Price, Kit of reference tests for cognitive factors. Educational Testing Service, 1963
work page 1963
-
[7]
Can a single human supervisor a swarm of 100 heterogeneous robots?
J. A. Adams, J. Hamell, and P. Walker, “Can a single human supervisor a swarm of 100 heterogeneous robots?” Journal of Field Robotics, DARPA OFFSET Special Issue, vol. 3, pp. 837–891, 2023. 39
work page 2023
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.