{"id":"737c9c8d-7d7f-42bb-8c54-2c665962d550","arxiv_id":"2411.18750","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Phase I baseline data from six Wing pilots shows no significant workload increase when active UAS count or nest count increases, with high situation awareness across nominal, DAA, and weather tasks.","lead":"This report evaluates a baseline human factors study of Wing delivery drone pilots supervising 10 to 100 active UAS. The finding is that doubling or tripling the number of drones and nests did not increase pilot workload or hurt situation awareness.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"24:High was always the second session, so UAS-count effects are confounded with practice; a non-significant result at n=3 per low arm does not support 'debunking' the fanout theory.","rationale":"The reader correctly identified both the small-sample/power issue and the session-order confound. I treat the session-order confound as the single most load-bearing concern because it is a design property that additional participants alone would not fix: even a large study would still compare a first-session X:Low condition with a second-session 24:High condition. The paper's own data contain direct evidence of a practice/novelty effect (declining first-session workload across Nominal tasks, and lower subjective workload and higher SA in the second session), and the authors explicitly decline inferential tests on in-situ workload for insufficient power in Section 3.3.4. The week-to-week active-UAS discrepancy is also unexplained and further weakens the manipulation, but it is secondary. The descriptive baseline data are plausible and useful, so I do not recommend moving to REJECT; the reader's CONDITIONAL verdict remains appropriate, with the condition being that the broad 'debunk' claim be either softened or supported by a counterbalanced design with equivalence testing.","tokens_in":30738,"tokens_out":7777,"duration_ms":76583,"concrete_test":"Counterbalanced replication: run six new PICs through the identical protocol with 24:High as the first session and one X:Low condition as the second (the reverse of the original order), using the same simulator, wearable sensors, and workload-scoring algorithm. Pre-register an equivalence margin, e.g., a mean overall workload difference of at most ±5 on the 0-100 scale. If first-session 24:High workload is materially higher than the original second-session 24:High data, or statistically equivalent to the original first-session X:Low values, the session-order confound is confirmed and the debunk claim fails; if first-session 24:High matches the original second-session 24:High mean, the order confound is unlikely to explain the null result.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the causal null claim in the abstract and executive summary, but the comparison supporting it is confounded: all six pilots completed 24:High in their second session and either 10:Low or 24:Low in their first (Section 2.1, Figure 2). Thus 'no workload difference' cannot be attributed to UAS count rather than practice and familiarization. The paper itself attributes improved SA and lower subjective workload in 24:High to learning effects (Sections 3.3.4 and 3.5), and Table 12 shows first-session workload declining across Nominal #1 to Nominal #3 (10:Low: 31.65 to 25.25; 24:Low: 34.22 to 24.53) — the same temporal pattern that would depress any second-session condition. The 24:High condition also differed by data-collection week (mean 90.3 vs 62.3 active UAS, Table 6) for reasons the authors state they could not identify, and each X:Low arm had only n=3 participants. A non-significant ANOVA at this sample size does not license 'debunk'; at most the study is an inconclusive baseline. The descriptive attention, SA, and interaction results remain useful, but the broad causal claim is not supported by the design.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a Phase I human-factors evaluation with six Wing-trained PICs in a high-fidelity simulated delivery-UAS environment. Three conditions were compared (10:Low, 24:Low, and 24:High), manipulating the number of nests and the number of active UAS, with nominal, DAA (single and double crewed aircraft), and adverse-weather tasks. The study collected objective physiological workload estimates from OSU's multi-dimensional workload algorithm, eye-tracking fixations, display-interaction logs, subjective in-situ workload, and SA probes. The results show overall workload in the normal-to-underload range, SA above 92% correct across trials, and no statistically reliable overall workload differences among conditions. The abstract and executive summary interpret these findings as debunking the theory that increasing the number of UAS degrades pilot performance.","tokens_in":30996,"tokens_out":5117,"duration_ms":47995,"significance":"If treated as a descriptive baseline, the dataset is valuable: it uses experienced, Part 107-certified PICs, Wing's actual delivery interface in a high-fidelity simulator, and a rich set of objective measures (physiological workload, eye tracking, interaction coding, SA probes). The reporting of descriptive statistics and the detailed AOI and interaction analyses are careful and will inform Phase II and regulatory conversations. However, the strong causal null claim ('debunk') is not supported by the design. The 24:High condition was always run second, the X:Low cells have n=3, the active-UAS count varied by week, and the large-df ANOVAs are pseudo-replicated. The paper's contribution is therefore best evaluated as a baseline study with descriptive findings, not as a test of the fanout theory.","major_comments":[{"comment":"The central claim that increasing the number of active UAS is not detrimental is confounded by session order. All six pilots completed 24:High in the second session, after completing either 10:Low or 24:Low in the first session (Section 2.1, Figure 2). The paper itself attributes the lower subjective workload and improved SA in 24:High to learning effects (Sections 3.3.4 and 3.5), and Table 12 shows first-session workload declining across Nominal #1 to Nominal #3 within the X:Low trials. A non-significant difference between the first-session X:Low and second-session 24:High conditions cannot be attributed to the number of UAS rather than to practice and familiarization. The wording 'debunk' is also at odds with the paper's stated hypothesis that these factors were predicted to have little impact (Section 1). This claim should be removed or substantially qualified, with the study presented as an inconclusive baseline.","section":"2.1, Figure 2; Executive Summary; Section 1"},{"comment":"The ANOVAs for the active-UAS and nest comparisons are reported with very high degrees of freedom (e.g., F(531, 2124) and F(531, 2655) in Sections 3.3.3.2 and 3.3.3.3), implying that time points within each 10-minute task were treated as repeated measures. With only three pilots per X:Low condition and six in 24:High, the participant-level error term is not properly represented; these tests are pseudo-replicated and cannot support generalizable null conclusions. The a-priori power analysis in Section 2.4 (n=6, f=0.25) is not applicable to the between-subject comparison of 10:Low versus 24:Low, each with n=3. The non-significant results in Section 3.3.2 should be reported as descriptive evidence only, and the authors should either fit mixed models with participant as a random effect or explicitly refrain from inferential claims about the number of UAS.","section":"3.3.3.2, 3.3.3.3; 2.4"},{"comment":"The 24:High manipulation was not internally consistent. The number of active UAS differed substantially by week (mean 90.33 in week 1 vs. 62.33 in week 2, Table 6), and the authors state they could not identify the cause. This means the 24:High condition was not a controlled 'high' UAS condition, and the comparison with X:Low aggregates across two different active-UAS levels, as well as different weather and DAA scripts. The abstract's claim that the results 'debunk' the theory is therefore unsupported even if the statistical issues in the previous comment were resolved. At minimum, the paper should restrict conclusions to the specific UAS ranges observed and acknowledge that the week-to-week variability limits the manipulation.","section":"3.1, Table 6; Abstract"}],"minor_comments":[{"comment":"The text contains a typo: 'ANOV A' should be 'ANOVA'. A global pass for such mechanical errors is needed.","section":"Section 3"},{"comment":"There is a grammatical error: 'and the a Shure microphone is worn as a headset' should read 'and a Shure microphone is worn as a headset.'","section":"Section 2.2"},{"comment":"The sentence 'There was virtually no difference between and withing the remaining component results' contains a typo ('withing' should be 'within').","section":"Section 3.3.4"},{"comment":"The sentence 'A 322% increase in fixation durations on AOI 2,2 existed as compared to the Nominal #1 task, with a 202% increase for the adjacent 2,2 AOI' is confusing because the adjacent AOI is also identified as 2,2; please clarify which AOI was intended.","section":"Section 3.2.3.1"},{"comment":"The table of contents lists two entries for Table 15 with slightly different titles ('The normalized in situ overall workload results descriptive statistics by trial' and 'The normalized in situ overall workload descriptive statistics by trial'), but only one Table 15 appears in the text. Please harmonize the numbering and titles.","section":"Table 15 (TOC)"},{"comment":"The derivation of the logarithmic workload model is clear, but the manuscript should more explicitly state that the fitted rate r=0.5 is based on prior work and is not validated against the Phase I data; this would avoid any impression that the model is used as empirical evidence.","section":"Section 2.5.1, Eq. (3)"}],"recommendation":"major_revision","confidential_remarks":"The 'debunk' phrasing in the abstract and executive summary is likely to attract media attention, but it is not supported by the design. I have framed this as a major comment rather than reject, because the underlying baseline data are useful and a careful revision that reframes the study as descriptive and explicitly acknowledges the session-order and week confounds would make an acceptable contribution. I recommend that the editor require the authors to either provide a participant-level mixed-model analysis or remove all inferential claims about the number of UAS. No concerns about novelty or citation practices."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a carefully run baseline study with unusually realistic equipment and real Wing PICs, and the descriptive results are worth having. But the abstract's 'debunk' claim is not supported by the design, and the paper needs to say so plainly.\n\nThe genuinely new thing here is the data: six trained Wing pilots using Wing's own delivery simulator and interface, with physiological workload estimates, eye tracking, SA probes, and coded interactions. That is a much richer environment than the simplified multi-UAS studies usually reported. The workload estimation and AOI coding appear careful, and the descriptive patterns (ADS-B dominant, Weather display during weather task, interactions mostly information-gathering) are coherent and internally consistent. The IMPRINT Pro predictions are clearly labeled as model-based, and the model is not used as evidence for the empirical claims.\n\nThe soft spot is the causal null claim. All six pilots flew 24:High in the second session, so the comparison against the X:Low conditions is confounded with practice and familiarization. The paper itself attributes improved SA and lower subjective workload in 24:High to learning effects, and Table 12 shows first-session workload declining across the Nominal tasks, the same temporal pattern that would depress a second session. With n=3 per X:Low arm, a non-significant ANOVA is not strong evidence for 'no effect.' The unexplained week-to-week difference in active UAS for 24:High (90 vs 62) adds another unplanned variable. The high-degrees-of-freedom repeated-measures ANOVAs (e.g., F(531,2655)) should be treated as secondary; the effective independent sample is tiny.\n\nNone of this kills the paper's value as a baseline. The descriptive results stand, and the authors are more careful in the body than in the abstract. But the 'debunks the traditional theory' framing needs to go, or be replaced by a statement that the study was not designed to test that theory at scale and that the null result is inconclusive.\n\nThe paper deserves a real referee — the data are rare, the method detail is useful, and Phase II will build on it. I'd send it out, and ask the authors to reframe the conclusion, explicitly address the session-order confound, and report the small-n caveat in the abstract itself.","headline":"Good baseline data from a real system; the abstract's 'debunk' claim outruns the design.","tokens_in":31533,"tokens_out":1809,"would_cite":true,"duration_ms":17404,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Raising the number of supervised delivery UAS from about 30 to more than 75 did not raise pilot workload in simulator trials.","keywords":["uncrewed aircraft systems","pilot workload","situation awareness","supervisory control","multi-UAS operations","delivery drones","eye tracking","workload modeling"],"falsifier":"Rerun the comparison with condition order counterbalanced and at least 10–20 pilots per condition; if mean physiological workload rises reliably above the low-UAS baseline once active UAS exceed roughly 75, the central claim fails.","tokens_in":30499,"feed_emoji":"📡","tokens_out":6160,"duration_ms":56169,"temperature":0.7,"pith_summary":"The paper reports a simulator study of professional delivery-drone pilots supervising large numbers of highly autonomous uncrewed aircraft systems (UAS). Across 60-minute trials, the authors varied the number of nests (10 vs. 24) and the number of active UAS (roughly 25–30 vs. 75–90), along with scripted crewed-aircraft encounters and an adverse-weather update. They claim that none of these manipulations produced a statistically reliable change in overall workload, that situation awareness stayed above 92 percent correct, and that pilots' visual attention and interface interactions remained consistent with their job duties. The paper presents these results as baseline evidence against the traditional theory that increasing the number of UAS degrades pilot performance.","feed_headline":"More delivery drones, same pilot workload in simulator trial","feed_subtitle":"With active UAS near 30 versus above 75, workload held near 29 on a 0-100 scale; SA stayed above 92%.","key_machinery":"The evaluation's central object is a logarithmic workload model, $w = a(1 + r \\ln n)$, which replaces the linear per-UAS workload assumption used in conventional modeling tools. Here $n$ is the number of active UAS, $a$ is the workload for a single UAS, and $r$ is a fitted logarithmic rate set to 0.5 for the delivery use case. This model carries the argument by predicting diminishing workload growth as UAS count increases, and the empirical result—flat physiological workload across UAS counts—is interpreted as confirming that sublinear relationship.","core_discovery":"The paper's central claim is that a single pilot in command can supervise substantially more autonomous delivery UAS without measurable performance cost, at least within the tested range. Increasing the number of active UAS by roughly 2.5 times left physiological workload estimates essentially flat, with overall means near 29 on a 0–100 scale and never approaching the overload threshold of 60; situation-awareness accuracy stayed at or above 92 percent. The authors argue this contradicts the fan-out theory developed for ground robots in the early 2010s, and they conclude that other factors—such as area size, UAS density, and geographical diversity—deserve more attention than raw UAS count.","pith_inferences":["If the flat-workload result reproduces, the binding constraint on how many delivery drones one pilot can supervise is likely airspace or interface complexity rather than raw UAS count.","The logarithmic model predicts a gradual eventual rise in workload as UAS count keeps increasing, so the cleanest extension is to probe much larger fleet sizes and locate that crossover.","Because every pilot flew the high-UAS condition second, a counterbalanced or between-subjects replication is the direct way to separate learning effects from any real fan-out cost."],"forward_implications":["Operators could treat the tested range (roughly 25–30 to 75–90 active UAS) as one where pilot workload does not automatically cap fleet size.","The results provide a baseline for the planned follow-up phase, which shifts attention to area size, UAS density, and geographical diversity as the factors likely to matter.","Interface and training emphasis should remain on monitoring and exception response rather than per-UAS control, since the ADS-B display absorbed most visual attention and interactions.","Workload modeling for highly autonomous multi-UAS operations should use sublinear growth in workload with UAS count rather than a linear per-UAS cost."],"supporting_citations":[{"why":"Field work with a swarm of 100 heterogeneous robots that the authors use to justify replacing the linear workload model with a logarithmic one.","marker":"[7]"},{"why":"The project final report whose insights are leveraged to derive the workload model for the delivery UAS use case.","marker":"[4]"},{"why":"The literature review that outlines existing multi-UAS results and the traditional assumptions the evaluation is designed to test.","marker":"[2]"},{"why":"A review of participant characteristics in multi-unmanned-vehicle studies, cited to support the claim that existing data lack ecological validity.","marker":"[1]"},{"why":"Prior measures of attention in multi-vehicle supervision, cited to motivate the objective eye-tracking and interaction metrics used here.","marker":"[3]"}],"fun_headline_variants":["UAS count up 2.5x, pilot workload flat: study","Debunking fan-out theory: more drones, same pilot load","Simulator trial: 75+ UAS, pilot SA stays above 92%","Pilot workload holds near 29 as UAS doubles: OSU-Wing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes that six professional pilots completing two 60-minute sessions, with the high-UAS condition always second, provide enough power and comparability to detect a real workload effect of UAS count.","fun_headline_variants_meta":{"raw":{"variants":["UAS count up 2.5x, pilot workload flat: study","Debunking fan-out theory: more drones, same pilot load","Simulator trial: 75+ UAS, pilot SA stays above 92%","Pilot workload holds near 29 as UAS doubles: OSU-Wing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000161,"raw_usage":{"total_tokens":1190,"prompt_tokens":854,"completion_tokens":336,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":252}},"tokens_in":470,"tokens_out":336,"duration_ms":4285,"temperature":1.0,"reasoning_tokens":252,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:54:17.570201+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the comparison with condition order counterbalanced and at least 10–20 pilots per condition; if mean physiological workload rises reliably above the low-UAS baseline once active UAS exceed roughly 75, the central claim fails.","supporting_citations":[{"cited_title":"Can a single human supervisor a swarm of 100 heterogeneous robots?","cited_arxiv_id":null,"evidence_quote":"Field work with a swarm of 100 heterogeneous robots that the authors use to justify replacing the linear workload model with a logarithmic one."},{"cited_title":"A26 A11L.UA V.74 : Establish pilot proficiency requirements multi- UAS components, final report,","cited_arxiv_id":null,"evidence_quote":"The project final report whose insights are leveraged to derive the workload model for the delivery UAS use case."},{"cited_title":"A26: Establish pilot proficiency requirements multi-uas components, literature review,","cited_arxiv_id":null,"evidence_quote":"The literature review that outlines existing multi-UAS results and the traditional assumptions the evaluation is designed to test."},{"cited_title":"Participant characteristics in human-in-the-loop studies with multiple unmanned vehicles including aircraft,","cited_arxiv_id":null,"evidence_quote":"A review of participant characteristics in multi-unmanned-vehicle studies, cited to support the claim that existing data lack ecological validity."},{"cited_title":"Measures of attention in autonomous and semi-autonomous multi-vehicle supervision,","cited_arxiv_id":null,"evidence_quote":"Prior measures of attention in multi-vehicle supervision, cited to motivate the objective eye-tracking and interaction metrics used here."}],"review_version":1}