Pith. sign in

REVIEW 3 major objections 6 minor 7 references

The Flight Physics Concept Inventory: Development of a research-based assessment instrument to enhance learning and teaching

T0 review · 3 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper introduces FliP-CoIn, the first research-based concept inventory for flight physics, developed bilingually, and reports internal-consistency reliability of .73.

desk verdict A transparent, genuinely useful instrument-development report whose reliability claim is somewhat over-sold by the headline alpha, but which deserves a serious referee. read the letter →

arxiv 2504.16975 v1 pith:IA4RYROI submitted 2025-04-23 physics.ed-ph physics.data-anphysics.flu-dyn

classification physics.ed-phphysics.data-anphysics.flu-dyn
keywords flightphysicsconceptinventoryeducationresearchstudentmisconceptionsaerodynamicliftbilingualassessmentreliabilitycontentvalidity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to establish that FliP-CoIn, a new multiple-choice instrument built in parallel in English and German, is a valid and reliable way to measure students' conceptual understanding of flight physics. The motivating gap is that no concept inventory for flight physics existed, despite a rich tradition of such instruments in mechanics and other introductory topics. The author argues that the instrument can diagnose naive concepts about lift, drag, stall, center of mass, streamlines, and angle of attack, and can therefore help instructors evaluate teaching and students see their own learning gaps. The headline evidence is an internal-consistency coefficient of .73 across 274 complete surveys, supported by expert reviews, think-aloud and focus-group interviews, and answer-pattern analyses. If the claim holds, flight-physics classrooms gain a shared, low-stakes formative assessment comparable to what the Force Concept Inventory provided for mechanics.

What carries the argument

The load-bearing mechanism is the concept inventory itself: a set of 46 scored items in forced-choice and ranking formats, built from learner-generated distractors, expert feedback, and think-aloud evidence. Each item is designed around a known naive concept or an expert concept, with distractors that are plausible to students holding that concept. A distinctive feature is that the instrument does not treat its concept domains, such as lift, drag, stall, center of mass, streamlines, and angle of attack, as statistically independent factors; the author argues these concepts are entangled and therefore declines factor analysis. Instead, the validity argument leans on answer-pattern analysis: when a small number of theoretically predicted answer patterns accounts for the large majority of responses, the patterns are taken as evidence that the items are measuring existing mental models. Scoring is dichotomous per item, with a total score whose internal consistency is the reported reliability evidence.

What would settle it

Run a bifactor or item-response analysis on the 274 complete booklets: if the general factor accounts for less than half of the common variance, or if high scorers systematically miss expert-validated items, the reported .73 overstates the coherence of what the total score measures.

Watch

Extended reading notes

Core claim

The central discovery is that student answers to flight-physics items cluster into a small number of recognizable conceptual patterns rather than random choices. For example, six answer patterns cover 58% of responses on a drag-ranking question, where 213 patterns were possible; on the drag-body sorting question, 89% of responses fall into six patterns plus their inverses. The author interprets this as evidence that stable naive concepts exist in flight physics and that a carefully built forced-choice inventory can surface them. The instrument's reliability is reported as .73, with several deliberately retained items that correlate negatively with the total score because they tap aspects, such as vortex drag behind a half-sphere or lift magnitude in steady climb, that no other item covers. Content validity is argued from an iterative design process involving eleven experts, interviews, and repeated revision over more than four years.

Load-bearing premise

The reliability claim rests on treating the total score from 46 deliberately entangled items as a meaningful single measure, despite the inventory keeping items that correlate negatively with that total.

Editorial extensions

If this is right

  • Instructors can use FliP-CoIn before and after instruction to estimate conceptual learning gains in flight physics with a common metric.
  • Teachers can identify which naive concepts are prevalent in a given class, such as the idea that a pointy nose is always best for drag or that lift must equal weight in any steady flight, and address them directly.
  • Because the instrument was developed in English and German from the start, cross-cultural comparisons of student thinking about flight physics become possible without translation artifacts.
  • The answer-pattern results allow researchers to quantify the prevalence of specific mental models rather than relying only on overall scores.
  • Keeping the instrument low-stakes and formative preserves its validity as a diagnostic tool for improving teaching and learning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension is to test whether a short subscale built from items with positive item-total correlations yields more sensitive pre/post gains, while the full instrument remains the richer diagnostic.
  • The bilingual development model could be applied to other technical subjects where terminology differs across languages, reducing translation-induced difficulty differences from the outset.
  • If the answer-pattern regularities replicate in new institutions, FliP-CoIn could serve as a culture-sensitive probe of teaching quality, not just individual learning.
  • Keeping negatively correlating items suggests a deliberate trade-off: reliability is sacrificed to preserve content coverage, and future revisions could quantify how much each retained item improves validity.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The manuscript is a doctoral thesis describing the development and validation of the Flight Physics Concept Inventory (FliP-CoIn), a bilingual (English/German) multiple-choice concept inventory for flight physics, intended for formative assessment of students' conceptual understanding of aerodynamic lift, drag, stall, center of mass, streamlines, angle of attack, and flight experience. The author documents a design-based research process that includes expert reviews, think-aloud and focus-group interviews, layman video interviews, multiple item revisions, and three pilot datasets (DS1, DS2, DS3; N=543 booklets, with n=274 complete surveys used in the final reliability analysis). The principal quantitative claim is a Cronbach's alpha of .725 across 46 items, interpreted as internal reliability, supported by answer-pattern analyses of drag-ranking items as content-validity evidence. The thesis also discusses gender differences, negatively correlated items that were deliberately retained, and recommendations for future research including test-retest reliability and predictive validity.

Significance. If the instrument were fully validated, it would fill a genuine gap: no flight-physics concept inventory had previously existed, and the bilingual from-scratch development is a distinctive contribution. The thesis is commendably transparent about the item evolution process, the coding of paper-pencil data (100% double coding with quantified expected undetected error), the dummy-user checks of online survey encoding, and the rationale for retaining difficult items. These strengths make the developmental narrative useful to the physics education research community. However, the psychometric evidence for the central claim is weakened by data-dependent scoring decisions and by the interpretation of Cronbach's alpha in an explicitly multidimensional instrument. The development process and qualitative validity evidence are valuable; the quantitative reliability claim needs substantial reframing or additional supporting analyses.

major comments (3)
  1. [Section 4.3 (iteration rounds) and Fig. 21] The final scoring rules are data-dependent in a way that invalidates the reported reliability coefficient as an estimate for the target population. The third-level iteration rounds state that item #001b was rescored to only the three answer patterns '1342', '1232', and '1243' and that item #033 was rescored to only options B and C after inspecting high- and low-scorer behavior; item retention decisions were likewise based on item-total correlations computed on the same DS1-DS3 datasets. The Cronbach alpha of .725 reported in Fig. 21 is therefore a fit statistic for the calibration data rather than an out-of-sample estimate, and the manuscript should either provide a cross-validation or clearly relabel the value as descriptive of the development sample.
  2. [Sections 3.6, 4.3.1, and 5.5] The interpretation of Cronbach's alpha as internal-consistency reliability presupposes that the items are essentially tau-equivalent indicators of a single construct, but the thesis explicitly argues that the FliP-CoIn items are intentionally multidimensional and entangled and declines factor analysis for this reason (Section 3.6). The final scale retains three items with negative corrected item-total correlations (Q5.11 = -.322, Q11.1 = -.10, Q19 = -.33 in Fig. 21), and Section 5.5 states that these items reduce alpha from .76 to .73 while being kept for content validity. Under these conditions, alpha is only an index of average inter-item covariance of a deliberately heterogeneous composite; it cannot by itself support the claim that the total score is a meaningful, coherent measure of flight-physics conceptual understanding. Please add dimensionality-robust reliability evidence (e.g., omega total or a justified composite-reliability approach) or restrict the claim to 'internal consistency of the calibrated composite'.
  3. [Section 3.3 and Abstract] The thesis states in Section 3.3 that determining reliability/precision in the sense of score consistency across repeated measurements is 'another task for future research,' yet the abstract and concluding sections present the instrument as reliable on the basis of the single internal-consistency coefficient. Because the central claim is about validity and reliability of a diagnostic instrument, the conclusions should be limited to what the data support: internal consistency in the calibration sample, with test-retest reliability, predictive validity, and cross-sample stability explicitly listed as unestablished. This is a revise-and-resubmit level issue, not a cosmetic one, because it changes the strength of the claims made to instructors.
minor comments (6)
  1. [Section 4.4, Question #5 analysis] The statement that six out of 213 possible answer patterns would cover 0.073% of answers under random responding appears arithmetically incorrect (6/213 is about 2.8%). The qualitative conclusion that observed concentration (58%) exceeds chance remains, but the numerical claim should be corrected.
  2. [Section 3.3.1] The formula for Cronbach's alpha is garbled in the text (the displayed equation contains stray characters); please typeset it properly with defined symbols.
  3. [Section 4.3.1 and Fig. 21] The text alternates between 'a=.73', 'Cronbach Alpha a=.73', and the table value '.725'; please use one consistent notation (e.g., α = .725) and add a confidence interval.
  4. [Section 5.4, Table 4] Calling the consistency between QID001c_1 and QID001c_5 'test-retest consistency' is misleading because the two items are in the same administration with three intervening questions on the same page; this is an immediate consistency check, not a test-retest reliability estimate, and the label should be changed.
  5. [Section 4.1, Table 1 and Section 5.1] The 'other/divers' gender category contains obvious joke entries and autofill artifacts noted in the text; to avoid misleading readers, the table should either flag this category as invalid or exclude it from descriptive comparisons.
  6. [Section 5.3, Table 3 caption and text] The ordering statement 'Drag Bodies (sorted from low to high aerodynamic drag)' conflicts with the discussion of which bodies students rank correctly; clarify whether the table lists expert ordering or item order, and make the expert ranking explicit.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the reliability and validity claims rest on transparent, independently sourced evidence rather than on a derivation that reduces to its own inputs.

full rationale

This is a psychometric validation thesis rather than a derivational chain, so the circularity inquiry reduces to whether the validation logic smuggles in its own conclusion. It does not. The reported Cronbach alpha (a=.73) is presented as a descriptive internal-consistency statistic computed after an explicitly iterative item-selection process; the paper does not call this a prediction or an external confirmation. The author is also transparent that test-retest reliability/precision is deferred to future research (Section 3.3), so the alpha is not dressed up as a complete reliability argument. The content-validity case is built from independent sources: expert reviews, think-aloud interviews, focus groups, laymen interviews, literature review, and answer-pattern analyses (Sections 3.5, 3.2.1, 4.4). No load-bearing self-citation chain was found: self-citations such as Genz and Bresges (2017) for design-based research are accompanied by independent methodological citations (Easterday et al., 2014; Scott et al., 2020), and the claimed gap in flight-physics concept inventories is supported by literature and platform searches rather than by the author's own prior work. The skeptical concern about multidimensionality and retained negative item-total correlations is a legitimate construct-validity critique, but it is not circularity: computing Cronbach's alpha for a deliberately multidimensional, entangled scale is a psychometric weakness, not an equation that reduces to its own inputs. Likewise, using the same datasets for item selection and for the final alpha is an in-sample overfitting concern, not a self-definitional or fitted-parameter-renamed-as-prediction step, because the paper never claims the alpha independently validates the item-selection choices made on the same data. The central validity claim therefore retains independent content and is not forced by definition or by self-citation.

Assumptions & free parameters 4 free parameters · 5 assumptions · 1 invented entities

No physical free parameters appear because the paper is an instrument-development study. The central claim depends on several data-dependent scoring and item-selection choices, listed as free parameters, and on psychometric assumptions about construct representation, dimensionality, expert judgment, and missing data. The Model Merging Phenomena is an auxiliary theoretical construct from an accompanying publication with no independent validation here.

free parameters (4)
  • Correct answer patterns for item #001b = 1342, 1232, 1243
    Rescored in Section 4.3, third-level iteration; patterns 1222 and 1234 were initially considered qualitatively correct but excluded because high scorers almost never used them.
  • Correct options for item #033 = B and C
    Initially scored B, C, D, E as correct; rescored to B and C after expert discussion and item analysis, affecting the final reliability estimate.
  • Item-retention threshold = item-total correlation below 0.15 triggers removal discussion
    Low-performing items were discussed, removed, modified, or reintroduced based on this data-dependent criterion applied to the same datasets used for the final alpha.
  • Complete-case cutoff = 95% of scored items answered
    A survey booklet was considered complete when 95% of scored items contained 0 or 1; this filter determined the 274 valid surveys used in the reliability analysis.
assumptions (5)
  • domain assumption Forced-choice multiple-choice and ranking responses can be scored dichotomously as indicators of expert-like conceptual understanding.
    The instrument assumes that correct and incorrect selections reflect underlying concepts; Section 4.4 uses answer-pattern analysis to argue that discrete mental concepts exist, but the mapping from response to concept is not directly demonstrated.
  • domain assumption The construct 'flight physics conceptual understanding' is represented by the seven selected domains: lift, drag, stall, center of mass, streamlines, angle of attack, and flight experience.
    Domain sampling was narrowed by expert feedback and local laboratory context (Section 3.4.2), but representativeness of the broader construct is asserted rather than formally demonstrated.
  • ad hoc to paper Factor analysis would be uninformative because the concepts are entangled, so dimensionality of the total score was not tested.
    This premise in Section 3.6 and Section 5.5.1 justifies reporting Cronbach's alpha without construct or dimensionality evidence; it is load-bearing for interpreting the total score.
  • domain assumption Expert judgment is a valid source of content validity.
    Eleven experts were interviewed and commented on items (Section 3.5.1), following psychometric convention; this is a reasonable but unverified assumption about the quality of the expert panel and their judgments.
  • domain assumption The complete-case analysis is unbiased despite only 274 of 543 surveys being complete.
    Authors chose listwise deletion because most dropouts occurred early (Section 3.8.6); if missingness correlates with ability, the reliability estimate may be biased.
invented entities (1)
  • Model Merging Phenomena (MMP)
    purpose: Describes students merging two coexisting but obscure models into one, filling gaps with naive model aspects; introduced in accompanying publication 2 to explain answer patterns (Section 7.3).
    Presented as a data-inspired theory from FliP-CoIn surveys; no external falsifiable prediction is specified in this thesis, and it is not central to the instrument validation claim.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Flight Physics Concept Inventory: Development of a research-based assessment instrument to enhance learning and teaching." pith.science (2026). https://pith.science/paper/IA4RYROI

@misc{pith2026250416975,
  author       = {Pith},
  title        = {Pith review of: The Flight Physics Concept Inventory: Development of a research-based assessment instrument to enhance learning and teaching},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IA4RYROI}},
  note         = {Machine review of arXiv:2504.16975}
}
read the original abstract

This work frames the first three publications around the development of the Flight Physics Concept Inventory (FliP-CoIn), and elaborates on many aspects in more detail. FliP-CoIn is a multiple-choice conceptual assessment instrument for improving fluid dynamics learning and teaching. I give insights into why and how FliP-CoIn was developed and how it is best used for improving conceptual learning. Further, this work presents evidence for several dimensions of FliP-CoIn's validity and reliability. Finally, I discuss key insights from the development process, the data analysis, and give recommendations for future research. This is a pre print version of the following book: Florian Genz, The Flight Physics Concept Inventory, 2025, Springer Spektrum, published with permission of Springer Fachmedien Wiesbaden GmbH. The final authenticated version is available online at: http://doi.org/10.1007/978-3-658-47515-4 and https://link.springer.com/book/9783658475147

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

7 extracted references · 7 canonical work pages

  1. [2]

    Formular re-check Method for FliP-CoIn's DS1,2,3 scoring calculations and descriptive Data

    Questions with differences in the codebook comparison were re-evaluated again. Then all questions were checked again with a special attention on typical places where the mistakes found so far occurred. 3. All score calculation formulars were re-checked and re-calculated if differences were discovered. The method is also described in the following video na...

  2. [5]

    The error occurred due to a wrong cell reference and the assumption that the third distractor is always coded with “3”

    In DS2 the count for the third distractor of v_QID024 was corrected. The error occurred due to a wrong cell reference and the assumption that the third distractor is always coded with “3”. The (correct) $-symbol in the formular of Cell KF494 led to the error that miscounted the cell value. Formular =COUNTIF(KF$383:KF$489;$JL494) was changed to =COUNTIF(KF...

  3. [6]

    =SUM(ES516:ES517)

    In DS2 the NA count calculation of v_example_understood (cell AK494) was =COUNTIF(AK$382:AK$489;AI494) instead of =COUNTIF(AK$383:AK$489;AI494). No influence on the count or scoring. (è No influence on the reliability analysis). 7. In DS2 the N count for (cell ES521) was corrected from “=SUM(ES516:ES517)” to “=SUM(ES516:ES520)”. No influence on the raw co...

  4. [8]

    NA counts

    In DS3 several formulars of “NA counts” were only counting “0” and were adapted accordingly to also count “f”(e.g. for v_QID032, v_QID020, v_QID006b) (è No in-fluence on the reliability analysis). 9. In DS3 the cell references for v_example_understood were corrected. This statistic was not used so far. No influence on any data analyses therefore (è No inf...

  5. [10]

    Formulars were copied from DS2 and re-checked on plausibility and consistency

    In DS3 calculation for pts1c_ALLcorrect was not yet implemented. Formulars were copied from DS2 and re-checked on plausibility and consistency. (è No influence on the reliability analysis, since not used in the final itemset). • • 11. In DS3 calculation of n and NA for v_QID_006d was corrected. No influence on scoring and data analysis (è No influence on ...

  6. [13]

    repeated word error

    In DS3 the calculation for NA and “repeated word error” calculations (Q34b_pattern in Column LV) was corrected from “=IF(LS576="NA-77";"NA-77";…” to “=IF(LS576="NA";"NA";…”. This had no influence on scoring and the rest of the data analysis. Only on the ratio of NA to “repeated word error” answers (è No influence on the reliability analysis). 14. In DS3 c...

  7. [29]

    Wave Diagnostic Test (WDT) 31

    Electromagnetics Concept Inventory (EMCI) 30. Wave Diagnostic Test (WDT) 31. Mechanical Wave Conceptual Survey (MWCS) 32. Four-tier Geometrical Optics Test (FTGOT) 33. Mechanical Waves Conceptual Survey 2 (MWCS2) 34. Wave Concept Inventory (WCI) 35. Thermal Concept Evaluation (TCE) 36. Survey of Thermodynamic Processes and First and Second Laws (STPFaSL) ...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.