Pith. sign in

REVIEW 3 major objections 6 minor 7 references

OSU-Wing PIC Phase I Evaluation: Baseline Workload and Situation Awareness Results

T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Raising the number of supervised delivery UAS from about 30 to more than 75 did not raise pilot workload in simulator trials.

desk verdict Good baseline data from a real system; the abstract's 'debunk' claim outruns the design. read the letter →

arxiv 2411.18750 v1 pith:ZAXNOIR2 submitted 2024-11-27 cs.HC cs.RO

classification cs.HCcs.RO
keywords uncrewedaircraftsystemspilotworkloadsituationawarenesssupervisorycontrolmulti-UASoperationsdeliverydroneseyetrackingmodeling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper reports a simulator study of professional delivery-drone pilots supervising large numbers of highly autonomous uncrewed aircraft systems (UAS). Across 60-minute trials, the authors varied the number of nests (10 vs. 24) and the number of active UAS (roughly 25–30 vs. 75–90), along with scripted crewed-aircraft encounters and an adverse-weather update. They claim that none of these manipulations produced a statistically reliable change in overall workload, that situation awareness stayed above 92 percent correct, and that pilots' visual attention and interface interactions remained consistent with their job duties. The paper presents these results as baseline evidence against the traditional theory that increasing the number of UAS degrades pilot performance.

What carries the argument

The evaluation's central object is a logarithmic workload model, $w = a(1 + r \ln n)$, which replaces the linear per-UAS workload assumption used in conventional modeling tools. Here $n$ is the number of active UAS, $a$ is the workload for a single UAS, and $r$ is a fitted logarithmic rate set to 0.5 for the delivery use case. This model carries the argument by predicting diminishing workload growth as UAS count increases, and the empirical result—flat physiological workload across UAS counts—is interpreted as confirming that sublinear relationship.

What would settle it

Rerun the comparison with condition order counterbalanced and at least 10–20 pilots per condition; if mean physiological workload rises reliably above the low-UAS baseline once active UAS exceed roughly 75, the central claim fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that a single pilot in command can supervise substantially more autonomous delivery UAS without measurable performance cost, at least within the tested range. Increasing the number of active UAS by roughly 2.5 times left physiological workload estimates essentially flat, with overall means near 29 on a 0–100 scale and never approaching the overload threshold of 60; situation-awareness accuracy stayed at or above 92 percent. The authors argue this contradicts the fan-out theory developed for ground robots in the early 2010s, and they conclude that other factors—such as area size, UAS density, and geographical diversity—deserve more attention than raw UAS count.

Load-bearing premise

The evaluation assumes that six professional pilots completing two 60-minute sessions, with the high-UAS condition always second, provide enough power and comparability to detect a real workload effect of UAS count.

Editorial extensions

If this is right

  • Operators could treat the tested range (roughly 25–30 to 75–90 active UAS) as one where pilot workload does not automatically cap fleet size.
  • The results provide a baseline for the planned follow-up phase, which shifts attention to area size, UAS density, and geographical diversity as the factors likely to matter.
  • Interface and training emphasis should remain on monitoring and exception response rather than per-UAS control, since the ADS-B display absorbed most visual attention and interactions.
  • Workload modeling for highly autonomous multi-UAS operations should use sublinear growth in workload with UAS count rather than a linear per-UAS cost.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the flat-workload result reproduces, the binding constraint on how many delivery drones one pilot can supervise is likely airspace or interface complexity rather than raw UAS count.
  • The logarithmic model predicts a gradual eventual rise in workload as UAS count keeps increasing, so the cleanest extension is to probe much larger fleet sizes and locate that crossover.
  • Because every pilot flew the high-UAS condition second, a counterbalanced or between-subjects replication is the direct way to separate learning effects from any real fan-out cost.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper reports a Phase I human-factors evaluation with six Wing-trained PICs in a high-fidelity simulated delivery-UAS environment. Three conditions were compared (10:Low, 24:Low, and 24:High), manipulating the number of nests and the number of active UAS, with nominal, DAA (single and double crewed aircraft), and adverse-weather tasks. The study collected objective physiological workload estimates from OSU's multi-dimensional workload algorithm, eye-tracking fixations, display-interaction logs, subjective in-situ workload, and SA probes. The results show overall workload in the normal-to-underload range, SA above 92% correct across trials, and no statistically reliable overall workload differences among conditions. The abstract and executive summary interpret these findings as debunking the theory that increasing the number of UAS degrades pilot performance.

Significance. If treated as a descriptive baseline, the dataset is valuable: it uses experienced, Part 107-certified PICs, Wing's actual delivery interface in a high-fidelity simulator, and a rich set of objective measures (physiological workload, eye tracking, interaction coding, SA probes). The reporting of descriptive statistics and the detailed AOI and interaction analyses are careful and will inform Phase II and regulatory conversations. However, the strong causal null claim ('debunk') is not supported by the design. The 24:High condition was always run second, the X:Low cells have n=3, the active-UAS count varied by week, and the large-df ANOVAs are pseudo-replicated. The paper's contribution is therefore best evaluated as a baseline study with descriptive findings, not as a test of the fanout theory.

major comments (3)
  1. [2.1, Figure 2; Executive Summary; Section 1] The central claim that increasing the number of active UAS is not detrimental is confounded by session order. All six pilots completed 24:High in the second session, after completing either 10:Low or 24:Low in the first session (Section 2.1, Figure 2). The paper itself attributes the lower subjective workload and improved SA in 24:High to learning effects (Sections 3.3.4 and 3.5), and Table 12 shows first-session workload declining across Nominal #1 to Nominal #3 within the X:Low trials. A non-significant difference between the first-session X:Low and second-session 24:High conditions cannot be attributed to the number of UAS rather than to practice and familiarization. The wording 'debunk' is also at odds with the paper's stated hypothesis that these factors were predicted to have little impact (Section 1). This claim should be removed or substantially qualified, with the study presented as an inconclusive baseline.
  2. [3.3.3.2, 3.3.3.3; 2.4] The ANOVAs for the active-UAS and nest comparisons are reported with very high degrees of freedom (e.g., F(531, 2124) and F(531, 2655) in Sections 3.3.3.2 and 3.3.3.3), implying that time points within each 10-minute task were treated as repeated measures. With only three pilots per X:Low condition and six in 24:High, the participant-level error term is not properly represented; these tests are pseudo-replicated and cannot support generalizable null conclusions. The a-priori power analysis in Section 2.4 (n=6, f=0.25) is not applicable to the between-subject comparison of 10:Low versus 24:Low, each with n=3. The non-significant results in Section 3.3.2 should be reported as descriptive evidence only, and the authors should either fit mixed models with participant as a random effect or explicitly refrain from inferential claims about the number of UAS.
  3. [3.1, Table 6; Abstract] The 24:High manipulation was not internally consistent. The number of active UAS differed substantially by week (mean 90.33 in week 1 vs. 62.33 in week 2, Table 6), and the authors state they could not identify the cause. This means the 24:High condition was not a controlled 'high' UAS condition, and the comparison with X:Low aggregates across two different active-UAS levels, as well as different weather and DAA scripts. The abstract's claim that the results 'debunk' the theory is therefore unsupported even if the statistical issues in the previous comment were resolved. At minimum, the paper should restrict conclusions to the specific UAS ranges observed and acknowledge that the week-to-week variability limits the manipulation.
minor comments (6)
  1. [Section 3] The text contains a typo: 'ANOV A' should be 'ANOVA'. A global pass for such mechanical errors is needed.
  2. [Section 2.2] There is a grammatical error: 'and the a Shure microphone is worn as a headset' should read 'and a Shure microphone is worn as a headset.'
  3. [Section 3.3.4] The sentence 'There was virtually no difference between and withing the remaining component results' contains a typo ('withing' should be 'within').
  4. [Section 3.2.3.1] The sentence 'A 322% increase in fixation durations on AOI 2,2 existed as compared to the Nominal #1 task, with a 202% increase for the adjacent 2,2 AOI' is confusing because the adjacent AOI is also identified as 2,2; please clarify which AOI was intended.
  5. [Table 15 (TOC)] The table of contents lists two entries for Table 15 with slightly different titles ('The normalized in situ overall workload results descriptive statistics by trial' and 'The normalized in situ overall workload descriptive statistics by trial'), but only one Table 15 appears in the text. Please harmonize the numbering and titles.
  6. [Section 2.5.1, Eq. (3)] The derivation of the logarithmic workload model is clear, but the manuscript should more explicitly state that the fitted rate r=0.5 is based on prior work and is not validated against the Phase I data; this would avoid any impression that the model is used as empirical evidence.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the workload and situation-awareness results are empirical outcomes, and the IMPRINT model is not used as evidence for the main conclusion.

full rationale

The paper's central claim that increasing the number of UAS and nests did not degrade pilot performance rests on measured physiological workload estimates, subjective workload ratings, interaction counts, eye-tracking data, and SA probe accuracy. These are data collected from the evaluation, not quantities defined in terms of the conclusion. The only potentially self-referential input is the IMPRINT Pro workload model, whose logarithmic rate r=0.5 is set from the authors' prior OFFSET and ASSURE work (Section 2.5.1), but the manuscript explicitly does not use that model's predictions as evidence for the empirical non-effect; Section 4 states that the model's predicted workload was substantially higher than the measured results and attributes this to modeling limitations. The OSU workload estimation algorithm is a measurement method rather than a fitted predictor of the condition effect, and no equation in the paper reduces the empirical outcome to an input. The SA scoring may be lenient in nominal tasks, but that is a construct-validity concern rather than a circular derivation. The confounded session order and small per-condition sample are experimental-validity risks, not circularity. No circular step was found, so the score is 0.

Assumptions & free parameters 2 free parameters · 4 assumptions · 1 invented entities

The central empirical claims do not depend on fitted constants; the main free parameter r=0.5 is only used for the IMPRINT Pro-model predictions, which are presented separately as predictions and are not used as evidence for the empirical workload conclusion. The key assumptions are the validity of the sensor-based workload algorithm, the representativeness of the simulation, and the adequacy of the small sample. No new physical entities are introduced.

free parameters (2)
  • IMPRINT Pro logarithmic rate r = 0.5
    Equation 3 defines workload as w = a(1 + r ln n). Section 2.5.1 states 'the logarithmic rate (r) for the IMPRINT Pro workload model was set to 0.5' based on prior OFFSET and ASSURE efforts. This value is fitted/chosen and is used to generate IMPRINT Pro predictions, though it is not used in the empirical debunking claim.
  • PIC ratio simulator parameters = 45 (10:Low), 35 (24:Low), 200 (24:High)
    These are manually chosen simulator settings intended to achieve target active UAS ranges of 25-30 or 80-100. They affect the number of active UAS, but are manipulated independent variables rather than fitted to the outcome.
assumptions (4)
  • domain assumption Workload can be modeled as w = a(1 + r ln n) for multi-UAS supervision.
    Section 2.5.1 introduces a logarithmic workload model to replace IMPRINT Pro's linear model. This is an assumed functional form, motivated by visual search efficiencies but not derived or validated against the human data in this paper.
  • ad hoc to paper The physiological sensor streams, when processed through OSU's multi-dimensional workload estimation algorithm, produce a valid scale of workload with the stated underload/normal/overload thresholds.
    Section 2.2 cites the algorithm without providing a validation reference or formal derivation, and the thresholds (<20, 20-59, >=60) are presented without external calibration evidence in this paper.
  • domain assumption The 60-minute simulated trials with scripted encounters and a static-slide weather display adequately represent real nominal operations and unexpected events for workload and SA measurement.
    The experimental setup replaces the real weather tool with a Google Slides METAR and uses scripted ADS-B traffic. The assumption that this preserves the workload-relevant aspects of the task is structural to the study and acknowledged only implicitly.
  • domain assumption The six participants, randomly split into 10:Low and 24:Low for session 1, provide enough power to detect a meaningful fanout effect.
    Section 2.4 reports a G*Power estimate, but the actual ANOVA analyses rely on repeated time-slices from three or six individuals, which is a small n assumption for the 'no effect' conclusion.
invented entities (1)
  • Adapted PIC Reference Sheet and Procedures Manual and scripted aircraft injection events
    purpose: Experimental materials created for this evaluation to standardize the PICs' information access and to generate DAA encounters.
    They are study scaffolding, not newly posited physical entities. They do not constitute a graviton-style invented entity, but are included for completeness of the ledger.

how reviews work

0 comments
Cite this review

Pith. "Pith review of OSU-Wing PIC Phase I Evaluation: Baseline Workload and Situation Awareness Results." pith.science (2026). https://pith.science/paper/ZAXNOIR2

@misc{pith2026241118750,
  author       = {Pith},
  title        = {Pith review of: OSU-Wing PIC Phase I Evaluation: Baseline Workload and Situation Awareness Results},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZAXNOIR2}},
  note         = {Machine review of arXiv:2411.18750}
}
read the original abstract

The common theory is that human pilot's performance degrades when responsible for an increased number of uncrewed aircraft systems (UAS). This theory was developed in the early 2010's for ground robots and not highly autonomous UAS. It has been shown that increasing autonomy can mitigate some performance impacts associated with increasing the number of UAS. Overall, the Oregon State University-Wing collaboration seeks to understand what factors negatively impact a pilot's ability to maintain responsibility and control over an assigned set of active UAS. The Phase I evaluation establishes baseline data focused on the number of UAS and the number of nests increase. This evaluation focuses on nominal operations as well as crewed aircraft encounters and adverse weather changes. The results demonstrate that the pilots were actively engaged and had very good situation awareness. Manipulation of the conditions did not result in any significant differences in overall workload. The overall results debunk the theory that increasing the number of UAS is detrimental to pilot's performance.

Figures

Figures reproduced from arXiv: 2411.18750 by the authors.

Figure 1
Figure 1. The monitors with the respective open windows for each trial session. The eye [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 2
Figure 2. An overview of the three trial conditions by session. [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. The task order and timing for each trial within all tasks. [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: The X:Low weather METARs. 2.2 Dependent Variables The dependent variables included the physiological data from the wearable sensors, the Wing simulator log files, the eye tracker, screen capture videos and camcorder videos, as well as the subjective in situ workload an…
Figure 5
Figure 5. Figure 5: The ADS-B display with the AOI numbered ( [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]
Figure 6
Figure 6. Figure 6: The r value analysis by the number of UAS for the nominal delivery use case. The IMPRINT Pro model also incorporates the human performance elements related to PICs responding to the in situ SA and workload rating probes during all tasks. The modeling of the probe respo…
Figure 7
Figure 7. Figure 7: The high-level developed Wing delivery use case IMPRINT Pro model. [PITH_FULL_IMAGE:figures/full_fig_p018_7.png]
Figure 8
Figure 8. Figure 8: The ADS-B display eye tracking fixations mapped to the AOIs by trial. Note [PITH_FULL_IMAGE:figures/full_fig_p023_8.png]
Figure 9
Figure 9. Figure 9: The DAA:X tasks fixations mapped to the AOIs by trial. The purple ovals [PITH_FULL_IMAGE:figures/full_fig_p027_9.png]
Figure 10
Figure 10. Figure 10: The mean overall workload estimates by tasks for (a) all trials, and (b) by week [PITH_FULL_IMAGE:figures/full_fig_p030_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

7 extracted references · 7 canonical work pages

  1. [1]

    Participant characteristics in human-in-the-loop studies with multiple unmanned vehicles including aircraft,

    E. J. Bass, R. Amey, J. Glavan, T. Read, N. Raghunath, C. A. Sanchez, K. Silas, T. Haritos, and J. A. Adams, “Participant characteristics in human-in-the-loop studies with multiple unmanned vehicles including aircraft,” in Human Factors and Ergonomics Annual Meeting, 2022

  2. [2]

    A26: Establish pilot proficiency requirements multi-uas components, literature review,

    E. Bass, J. A. Adams, C. A. Sanchez, and T. Haritos, “A26: Establish pilot proficiency requirements multi-uas components, literature review,” F AA’s ASSURE Center of Ex- cellence, Tech. Rep., February 2021

  3. [3]

    Measures of attention in autonomous and semi-autonomous multi-vehicle supervision,

    J. J. Glavan, E. J. Bass, C. A. Sanchez, T. Read, and J. A. Adams, “Measures of attention in autonomous and semi-autonomous multi-vehicle supervision,” in IEEE International Conference on Human-Machine Systems, 2022

  4. [4]

    A26 A11L.UA V.74 : Establish pilot proficiency requirements multi- UAS components, final report,

    J. A. Adams, P. Uriarte, C. A. Sanchez, T. Read, J. Glavan, E. J. Bass, T. Hari- tos, and K. Silas, “A26 A11L.UA V.74 : Establish pilot proficiency requirements multi- UAS components, final report,” F AA’s ASSURE Center of Excellence, Tech. Rep. A26 A11L.UA V.74, November 2022

  5. [5]

    The development of a short domain-general measure of working memory capacity,

    F. L. Oswald, S. T. McAbee, T. S. Redick, and D. Z. Hambrick, “The development of a short domain-general measure of working memory capacity,” Behavior research methods, vol. 47, pp. 1343–1355, 2015

  6. [6]

    E. T. Service, J. W. French, R. B. Ekstrom, and L. A. Price, Kit of reference tests for cognitive factors. Educational Testing Service, 1963

  7. [7]

    Can a single human supervisor a swarm of 100 heterogeneous robots?

    J. A. Adams, J. Hamell, and P. Walker, “Can a single human supervisor a swarm of 100 heterogeneous robots?” Journal of Field Robotics, DARPA OFFSET Special Issue, vol. 3, pp. 837–891, 2023. 39

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.