{"id":"fe321b73-3cc3-411a-b8e6-0834115a9008","arxiv_id":"2501.17028","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":8,"one_line_summary":"A semi-automated certification framework for DO-178C Level D ML systems is demonstrated on a YOLOv8 vehicle detector, producing a Moderate Assurance score of 74.7.","lead":"This paper proposes a semi-automated certification framework for low-criticality machine learning systems in aviation, combining structured classification, an Assurance Profile score, and human oversight. A YOLOv8 vehicle-detection case study yields a Moderate Assurance certification score of 74.7.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Weighted-average Assurance Profile cannot establish DO-178C Level D compliance: QA and CM scores fall below the paper's own insufficiency threshold, yet Section V-A claims compliance.","rationale":"The reader's REJECT is justified, but the reason should be sharpened. Arbitrary weights in Table I are a symptom; the deeper defect is that the Assurance Profile is a compensatory scoring model applied to a non-compensatory standard. DO-178C Level D compliance requires each applicable objective to be independently satisfied, whereas the paper's weighted average lets strong Development and V&V results mask weak QA and CM results. The paper's own per-process scores demonstrate this: QA is 52.5, below its own 'Insufficient' threshold, and CM is 61.5, 'Limited'. The conclusion in Section V-A that Air Sight 'meets the compliance criteria' is therefore not merely an issue of input calibration but a mismatch between the proposed evidence aggregation and the structure of the certification standard. I also credit the paper's genuine strengths: it explicitly scopes to static, offline-trained ML models; it builds on EASA guidance and prior certification-oriented work; it uses a concrete tool (Deepchecks) for drift checks; and it outlines recertification triggers. These contributions support a 'pre-certification readiness assessment' framing, but they do not bridge the gap to a DO-178C Level D compliance claim. A single conjunctive re-evaluation of Table I would settle the issue, and as the table stands the claim cannot be sustained. The reader's verdict should remain REJECT; no verdict change is needed.","tokens_in":8762,"tokens_out":5147,"duration_ms":52400,"concrete_test":"Re-run the certification decision using a conjunctive gate: apply the paper's own sigma thresholds to each process score in Table I and require all four processes to be at least at 'Moderate Assurance' (S >= 70) with no individual activity below 60. Under Table I, QA (52.5) and CM (61.5) fail, so Section V-A's 'meets compliance' claim would not follow from the authors' own numbers. This check distinguishes the paper's compensatory averaging from the objective-based, all-objectives-satisfied logic required by DO-178C Level D.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is not just that Table I weights and activity scores are hand-assigned, but that a weighted average can certify compliance. Section IV-C defines the final score as a weighted sum of process scores and maps S to confidence levels. In Table I, QA scores 52.5 and CM scores 61.5, which are below the paper's own 'Insufficient Assurance' (<60) and 'Limited Assurance' (60-70) boundaries. Yet the high Development and V&V scores pull the aggregate to 74.7 ('Moderate Assurance'), and Section V-A states: 'This initial certification assessment demonstrates that Air Sight meets the compliance criteria for DO-178C Level D criticality.' DO-178C is an objective-based standard: each applicable Level D objective must be satisfied; there is no provision for a high V&V score to compensate for an unsatisfied QA or configuration-management objective. The paper supplies no mapping from Assurance Profile activities to DO-178C Level D objectives and no argument that 'Moderate Assurance' is equivalent to passing those objectives. Thus, even if all weights in Table I were externally justified, the central claim would still fail because the scoring model is compensatory while the standard is conjunctive. The paper's own numbers make the compliance statement internally inconsistent with its rubric: two of four process areas are in the limited or insufficient range, yet the aggregate is presented as meeting compliance.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a semi-automated certification approach for low-criticality (DO-178C Level D) ML-enabled airborne systems. The approach consists of a multi-axial system classification C=(ccrit, caut, cmodel), three certification layers, and an Assurance Profile that aggregates manual and automated activity scores into a weighted process score and a final certification score S. The final score is mapped to a confidence level sigma(S). The approach is exercised on a case study, Air Sight, a YOLOv8-based vehicle detector for surveillance aircraft. The paper reports a final score of 74.7, labels it 'Moderate Assurance', and concludes in Section V-A that 'Air Sight meets the compliance criteria for DO-178C Level D criticality'. The paper also defines recertification triggers and reports a robustness evaluation with Gaussian noise and drift checks.","tokens_in":9097,"tokens_out":4102,"duration_ms":42178,"significance":"The paper is clearly relevant to an active problem: current airworthiness standards do not address ML-specific assurance artifacts such as data validation, drift monitoring, and model robustness. Its strengths are that it is explicitly scoped to Level D and to offline-trained, static models; it makes its assumptions explicit; it reports a real case study with YOLOv8 and Deepchecks; and it provides a public code repository link for the prototype. If the proposed framework could demonstrate a defensible link to DO-178C objectives, it would be a useful contribution to the ongoing discussion on ML certification. However, as written, the central claim is not supported: the Assurance Profile is a compensatory score computed from hand-assigned inputs, whereas DO-178C compliance is an objective-based, conjunctive determination. The paper therefore currently functions more as an illustrative position piece than as a validation of a certification method.","major_comments":[{"comment":"The certification score is computed as a weighted sum S = sum_j w_process_j * S_process_j, which is a compensatory model: high Development and V&V scores offset low QA and CM scores. In Table I, the QA score is 52.5 and the CM score is 61.5, which fall below the paper's own 'Insufficient Assurance' (<60) and 'Limited Assurance' (60-70) thresholds, respectively. Despite this, Section V-A states that 'Air Sight meets the compliance criteria for DO-178C Level D criticality'. DO-178C is an objective-based standard: each applicable objective must be satisfied, and there is no provision for a high verification score to compensate for an unsatisfied quality-assurance or configuration-management objective. The paper supplies no mapping from the Assurance Profile activities to the DO-178C Level D objectives, so the compliance statement is internally inconsistent with the paper's own rubric and unsupported by the cited standard.","section":"Section IV-C and Table I, Section V-A"},{"comment":"The weights w_act and w_process are introduced as 'a factor of the system's classification C and contextual factors', but no elicitation procedure, data source, sensitivity analysis, or independent criterion is given for their values. The activity scores in Table I are reported without confidence intervals, inter-rater agreement, or an audit trail showing how manual reviews and automated check pass rates were converted into the 0-100 scores. Since the final score 74.7 is an arithmetic consequence of these hand-assigned inputs, the resulting 'Moderate Assurance' verdict is not an independent measurement or prediction. Before the score can support a compliance determination, the paper should provide a documented scoring protocol, a sensitivity analysis, and ideally a comparison with a certification authority or with a system of known conformance status.","section":"Section IV-C, score equations and Table I"},{"comment":"The evaluation concludes that 'prediction drift was observed during testing, it remained within the acceptable threshold for Level D criticality systems (<30%)', but the only 30% threshold defined in Section V-B is a recertification trigger for 'a significant change (over 30%) in dataset distribution or model architecture'. The paper does not define an acceptable prediction-drift threshold for Level D, nor does it justify transferring the 30% value from dataset change to prediction drift. The conclusion that recertification is not needed is therefore based on an arbitrary threshold and should either be removed or supported by a rationale tied to operational requirements or an applicable standard.","section":"Section V-C versus Section V-B"}],"minor_comments":[{"comment":"The phrase 'confidence measure the ML component' is missing a preposition and should read 'confidence measure for the ML component'.","section":"Abstract"},{"comment":"The text contains a duplicated phrase: 'low-criticality corresponds to DO-178C Level D, corresponds to DO-178C Level D'. The repetition should be removed.","section":"Section I"},{"comment":"The summation notation 'Sprocess = nX i=1 wacti ·Sacti' uses an undefined upper limit n; please define the number of activities per process or use explicit enumeration.","section":"Section IV-C"},{"comment":"The Configuration Management process is labeled 'Total SCM Score' in the table but is consistently called CM elsewhere in the paper; the label should be made consistent.","section":"Table I"},{"comment":"There are several typographical errors, including 'uncertainity' in the introduction and 'appropraite' in Section IV-C5; the manuscript should be proofread for such issues.","section":"Section IV-C"}],"recommendation":"reject","confidential_remarks":"The central problem is not the absence of fine-tuning but the mismatch between the paper's claimed certification outcome and the method actually implemented: a weighted average cannot establish satisfaction of a conjunctive, objective-based standard. I would not rule out a future version of this work, but it would need a fundamentally different argument structure, including an objective-level traceability matrix and a non-compensatory treatment of failing assurance areas. As it stands, the Section V-A compliance claim is not defensible within the manuscript's current scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The three-axis classification (criticality, autonomy, model complexity) and the Assurance Profile score are genuinely new relative to the cited prior work, and the case study is concrete. The recertification triggers are sensible. This is a useful assembly of existing ideas—not a breakthrough, but it addresses a real gap in certifying low-criticality ML systems.\n\nThe soft spot is the central claim. The final score in Table I is a weighted sum of hand-assigned activity scores, and no procedure is given for deriving the weights. More damagingly, the paper's own rubric says QA (52.5) is insufficient and CM (61.5) is limited, yet the aggregate 74.7 is called Moderate and then interpreted as \"meets compliance criteria\" for DO-178C Level D. DO-178C is conjunctive: each applicable objective must be satisfied. There is no provision for a strong V&V score to offset a missing QA or CM objective. The paper never maps its activities to DO-178C Level D objectives, so even if the weights were externally justified, the compliance conclusion would not follow. The stress-test note lands.\n\nAlso missing: the code and data needed to reproduce the demonstration, and the reported model metrics lack error bars. These are minor compared to the compliance issue.\n\nCredit where earned: the authors are honest about the weaknesses in QA and CM, and they frame the work as preliminary. The framework is a plausible starting point for a certification assessment, just not a compliance proof.\n\nMy recommendation: this deserves a serious referee because the problem is real and the approach is worth debating. But it needs major revision—either drop the compliance claim or map each activity to specific Level D objectives and show all are satisfied. I would send it to review with an expectation of major revisions, not desk reject.","headline":"Useful framework, but the compliance claim doesn't survive the paper's own rubric.","tokens_in":9572,"tokens_out":2340,"would_cite":false,"duration_ms":22471,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A weighted Assurance Profile can certify low-criticality ML systems for DO-178C Level D.","keywords":["machine learning certification","DO-178C","low criticality airborne systems","assurance profile","semi-automated certification","object detection","YOLOv8","ML assurance"],"falsifier":"Recompute the final Assurance Profile with a different but plausible Quality Assurance process weight, say 0.35 instead of 0.20; the total falls below 70, flipping the verdict from Moderate to Limited Assurance and contradicting the paper's claim that Air Sight meets DO-178C Level D compliance.","tokens_in":8576,"feed_emoji":"✈️","tokens_out":5346,"duration_ms":46435,"temperature":0.7,"pith_summary":"This paper proposes a semi-automated route to certifying machine-learning components in low-criticality airborne systems, targeting DO-178C Level D where failures cause only minor conditions. The central idea is to classify the ML system along criticality, autonomy, and model complexity, then consolidate automated checks and human reviews into a single Assurance Profile score on a 0-100 scale. The authors demonstrate the approach on Air Sight, a YOLOv8-based object detector for reconnaissance and surveillance aircraft, and report a final score of 74.7, which their benchmark labels Moderate Assurance. They conclude that Air Sight meets Level D compliance and needs only periodic drift checks, not full recertification, during its operational life.","feed_headline":"Weighted scorecard certifies ML aircraft detector at Level D","feed_subtitle":"A semi-automated Assurance Profile turns checks and reviews into one number, 74.7, for DO-178C compliance.","key_machinery":"The central object is the Assurance Profile: a structured scorecard, described by the authors as similar to a nutrition label, that consolidates all certification evidence into one number. Activity scores $S_{\\mathrm{act}}$ are aggregated into process scores by weighted sums, and the four process scores are aggregated into the final certification score $S$. The weights are chosen as a function of the system classification $C = \\langle c_{\\mathrm{crit}}, c_{\\mathrm{aut}}, c_{\\mathrm{model}} \\rangle$ and contextual factors. The classification also selects the depth of validation: Air Sight's complex model (a YOLOv8 neural network) receives the most extensive validation level $V_3$, while Level D criticality and 2A autonomy keep human oversight in the loop.","core_discovery":"The paper's core claim is that certification readiness for a low-criticality ML component can be expressed as a single weighted score derived from four processes: Development, Verification & Validation, Quality Assurance, and Configuration Management. Each process is scored through activities that mix automated checks, semi-automated tests, and manual reviews, and the scores are combined using weights that depend on the system's classification $C = \\langle c_{\\mathrm{crit}}, c_{\\mathrm{aut}}, c_{\\mathrm{model}} \\rangle$. For Air Sight, classified $\\langle D, 2A, 3 \\rangle$, the resulting score is $S = 74.7$, corresponding to Moderate Assurance. The paper states that this result demonstrates that Air Sight meets the compliance criteria for DO-178C Level D and that, for its classification, full recertification is unnecessary within a typical operational lifecycle; targeted drift checks suffice.","pith_inferences":["The weakest link in the chain is the provenance of the Assurance Profile weights: until a derivation or calibration procedure is published, the framework can produce a different certification verdict for the same system under a different but equally plausible set of hand-assigned numbers.","A natural next test is to apply the same scorecard to several Level D ML systems and check whether the resulting score separates systems that later experience operational failures from those that do not.","The 30% drift threshold used as a recertification trigger is presented as acceptable for Level D; tying it to a formal risk model would make recertification decisions auditable across applications."],"forward_implications":["Certification of a Level D ML component can reduce to producing a documented Assurance Profile, with automated checks doing most of the verification work and manual reviews reserved for integration, usability, and uncertainty handling.","Once certified, the system can remain compliant through targeted drift checks rather than full re-audits, as long as drift stays below the stated thresholds such as the 30% dataset-shift trigger.","Weak scores in Quality Assurance and Configuration Management do not block certification but identify where post-deployment monitoring and version-control discipline must improve.","The same classification and scorecard structure can be applied to other low-criticality ML-enabled airborne functions with comparable autonomy and model complexity."],"supporting_citations":[{"why":"Defines DO-178C and the Level D criticality category that the certification approach targets.","marker":"[1]"},{"why":"EASA guidance supplies the autonomy levels used in the classification and documents gaps in existing standards for ML behavior.","marker":"[9]"},{"why":"Establishes the low-criticality MLS certification problem that this work builds on and extends.","marker":"[10]"},{"why":"The ML Test Score provides the idea of structured test cases and scoring that the Assurance Profile adapts.","marker":"[17]"},{"why":"YOLOv8 is the object detection model used in the Air Sight case study.","marker":"[11]"},{"why":"The military vehicles dataset is used to fine-tune and validate the Air Sight detector.","marker":"[30]"},{"why":"Supplies the automated drift-detection checks used in the evaluation.","marker":"[31]"}],"fun_headline_variants":["Weighted score 74.7 earns ML detector Level D clearance","YOLOv8 detector passes Level D with semi-auto certification","Semi-auto scorecard lets drift checks replace recertification","Airborne ML gets Level D via one weighted assurance score","Assurance Profile scores 74.7 for low-criticality aircraft ML"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The certification verdict rests entirely on Assurance Profile weights and activity scores that are assigned by hand without a documented derivation, so choosing different numbers would produce a different compliance outcome.","fun_headline_variants_meta":{"raw":{"variants":["Weighted score 74.7 earns ML detector Level D clearance","YOLOv8 detector passes Level D with semi-auto certification","Semi-auto scorecard lets drift checks replace recertification","Airborne ML gets Level D via one weighted assurance score","Assurance Profile scores 74.7 for low-criticality aircraft ML"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000227,"raw_usage":{"total_tokens":1442,"prompt_tokens":888,"completion_tokens":554,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":504,"completion_tokens_details":{"reasoning_tokens":466}},"tokens_in":504,"tokens_out":554,"duration_ms":5972,"temperature":1.0,"reasoning_tokens":466,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T05:01:19.926868+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the final Assurance Profile with a different but plausible Quality Assurance process weight, say 0.35 instead of 0.20; the total falls below 70, flipping the verdict from Moderate to Limited Assurance and contradicting the paper's claim that Air Sight meets DO-178C Level D compliance.","supporting_citations":[{"cited_title":"DO-178C: Software Considerations in Airborne Systems and Equipment Certification,","cited_arxiv_id":null,"evidence_quote":"Defines DO-178C and the Level D criticality category that the certification approach targets."},{"cited_title":"Easa artificial intelligence (ai) concept paper issue 2: Guidance for level 1&2 machine learning applications,","cited_arxiv_id":null,"evidence_quote":"EASA guidance supplies the autonomy levels used in the classification and documents gaps in existing standards for ML behavior."},{"cited_title":"Toward certification of machine-learning systems for low criticality airborne applications,","cited_arxiv_id":null,"evidence_quote":"Establishes the low-criticality MLS certification problem that this work builds on and extends."},{"cited_title":"The ml test score: A rubric for ml production readiness and technical debt reduction,","cited_arxiv_id":null,"evidence_quote":"The ML Test Score provides the idea of structured test cases and scoring that the Assurance Profile adapts."},{"cited_title":"Jocher et al., “Yolov8.” https://github.com/ultralytics/ultralytics, 2023","cited_arxiv_id":null,"evidence_quote":"YOLOv8 is the object detection model used in the Air Sight case study."},{"cited_title":"Military vehicles object detection dataset","cited_arxiv_id":null,"evidence_quote":"The military vehicles dataset is used to fine-tune and validate the Air Sight detector."},{"cited_title":"Deepchecks: A library for testing and validating machine learning models and data,","cited_arxiv_id":null,"evidence_quote":"Supplies the automated drift-detection checks used in the evaluation."}],"review_version":1}