{"id":"1785808f-0028-4e63-b6d6-2406ce5195d0","arxiv_id":"2506.11525","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A step-by-step framework with five steps, eleven input factors, and seven desirability categories supports systematic assessment of process deviations.","lead":"This paper introduces a five-step framework that helps process analysts decide whether deviations from a process model are false alarms, justified exceptions, or genuinely positive or negative and worth acting on. It was built from a literature review and practitioner interviews, and tested with eight analysts who found it useful.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"P1 rests on an untested reliability assumption: the evaluation never checks whether analysts assign the same mutually exclusive category, and Section 6 concedes category ambiguity.","rationale":"I read the paper as a design-science contribution: the contribution is not a theorem but a procedural instrument, so the relevant burden is evidence that using the instrument yields stable, correct, actionable classifications. The framework is well structured, grounded in a documented literature review and interview/focus-group process, and the qualitative evaluation does provide some evidence for P2 (orientation, structured assessment, especially for novices). I credit that evidence. The load-bearing weakness is that P1—correctness and completeness of components, and mutual exclusivity of the seven categories—is not tested in the way the claim requires. Section 5.1's explicit decision not to compare participants' outcomes means that the reported convergence ('mostly converged') is anecdotal, and Section 6's admission of possible ambiguity directly conflicts with the 'mutually exclusive' characterization in the abstract. The missing aggregation rule for the four impact factors is a concrete instance: without a stated way to weigh, e.g., positive performance against negative compliance, different analysts can legitimately reach different categories from the same inputs, so the framework does not by itself determine the output. This is not a disagreement with the field's consensus; it is an internal gap between the claimed guarantee and the specified procedure. A reliability study with independent raters is the natural way to settle it. I would keep the reader's CONDITIONAL verdict: the paper makes a plausible contribution, but the core classification reliability claim should be demonstrated before acceptance as established.","tokens_in":13408,"tokens_out":4680,"duration_ms":43951,"concrete_test":"Run a randomized inter-rater reliability study: recruit two independent panels of at least five experienced process analysts, give each analyst the same 20–30 deviations from real or realistic event logs with full case context (identical to what the framework assumes), and have them apply the framework independently. Measure agreement on the seven output categories with Fleiss' kappa, and compare with a reference panel that adjudicates each deviation. Decision rule: if kappa is below 0.6 or more than about 10% of deviations receive different categories across raters, the framework does not reliably yield the claimed mutually exclusive, action-linked categories, so P1 remains unsupported. A secondary check can isolate Step 4 by presenting counterbalanced cases that trade off compliance against performance to see if analysts' category choices vary systematically.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is P1: the framework's five steps and eleven input factors are correct and complete, and they partition deviations into seven mutually exclusive categories. The evaluation does not actually test this. Section 5.1 states that participants' assignments were not compared: 'we did not evaluate whether participants arrive at a certain outcome' because interpretative assumptions made classification outcomes incomparable. The only P1 evidence is that eight evaluators did not spontaneously report missing factors or misalignments on four synthetic examples. That is an absence-of-complaint result, not a correctness or reproducibility result. Section 6 then concedes 'there might be ambiguity within output categories, meaning that, e.g., a deviation might be an exception in some cases whereas it is a negative deviation in other cases,' which undercuts the advertised mutual exclusivity. Step 4 (Evaluate impact) also has no aggregation rule for the four factors—compliance, outcome, performance, standardization—and the paper says trade-offs 'are typically discussed with the management of the process.' Consequently, the procedure underdetermines the final category whenever factors conflict, and the claimed action-linked categories are not fully determined by the framework. The central claim therefore depends on an unverified reliability assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a procedural framework for assessing the desirability of process deviations, developed through a literature review (32 papers) and qualitative interviews and a focus group. The framework consists of two analysis parts and five steps, using eleven input factors as knockout criteria to assign seven output categories (false alarm log, false alarm model, exception, neutral, positive, negative, reaction-inefficient), each with an action recommendation. The authors claim that the framework's components are correct and complete (P1) and that it is useful for process analysts (P2), and they evaluate it through a deviation assessment task with eight practitioners. The paper includes supplementary materials and explicitly acknowledges limitations in quantitative evaluation.","tokens_in":80,"tokens_out":3412,"duration_ms":74322,"significance":"If validated, the framework would fill a genuine gap: conformance checking identifies deviations but not their desirability. The work usefully synthesizes scattered literature and provides a structured, actionable procedure. The authors are transparent about their qualitative method and provide open supplementary materials. However, the current evidence for P1 is limited: there is no inter-rater reliability analysis, no gold-standard correctness test, and the completeness claim rests on the absence of reported missing factors from eight evaluators. The paper's honest acknowledgment of these limitations is a strength, but it also means the central claims are not yet fully supported.","major_comments":[{"comment":"Section 5.1 states that the evaluation did not examine whether participants arrive at a certain outcome because 'interpretative assumptions render classification outcomes incomparable.' Consequently, the paper provides no reliability evidence for the central claim that the framework categorizes deviations into mutually exclusive categories. The only support for P1 is the absence of reported gaps from eight practitioners, which is a weak basis for claiming correctness and completeness. At minimum, the authors should report a comparison of participants' final categories against the researchers' expected categories, or temper the P1 claim to a plausibility argument.","section":"Section 5.1"},{"comment":"The Evaluate impact step considers four input factors (compliance, outcome, performance, and standardization) but provides no prescriptive aggregation rule. The paper states in Section 6 that trade-offs among these factors 'are typically discussed with the management of the process.' This means the framework underdetermines the final desirability category whenever the factors conflict, as in the paper's own invoice example where time savings are weighed against compliance violations. The authors should either specify a decision rule or explicitly characterize the framework as an analytical aid rather than a deterministic classifier.","section":"Section 4.2, Evaluate impact"},{"comment":"The paper concedes 'there might be ambiguity within output categories, meaning that, e.g., a deviation might be an exception in some cases whereas it is a negative deviation in other cases.' This directly conflicts with the abstract's claim of 'mutually exclusive desirability categories.' The authors should clarify whether mutual exclusivity is guaranteed by the procedure or is merely a design goal, and discuss how ambiguity should be resolved in practice.","section":"Section 6, Application Challenges"}],"minor_comments":[{"comment":"The phrase 'afalse alarm (log)' in the running text should be corrected to 'a false alarm (log)'.","section":"Section 4.1"},{"comment":"Figure 2 is visually dense; a legend that clearly distinguishes input factors, assessment steps, output categories, and action recommendations would improve readability.","section":"Figure 2"},{"comment":"The statement that 'the experts mostly converged on similar assessments' would be more informative if the paper reported the level of agreement or disagreement, for example by listing the categories assigned by each participant to a sample deviation.","section":"Section 5.2"},{"comment":"The repeated phrase 'out of scope' used for several expert suggestions could be replaced with a more structured discussion of the boundary conditions of the framework.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a design-science/qualitative contribution; for a process mining venue, the evaluation is on the weaker side. The authors are honest about limitations, but the central claim requires either a more rigorous evaluation or a reframing. I would be willing to consider a revised version that clearly states what is and is not claimed and, if feasible, includes an inter-rater reliability check or a decision rule for conflicting factors."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper is worth a look. What's genuinely new is the procedural structure itself: prior work on deviation desirability offered categories (false alarms, exceptions, workarounds, severity), but not an ordered assessment procedure with knockout criteria and linked actions. The five-step framework with eleven input factors and seven output categories is a real assembly job, and the authors are transparent about how it was built: literature coding, development interviews, a focus group, then evaluation interviews with eight fresh practitioners. The 32-paper review is documented and supplementary materials are provided.\n\nThe soft spot is the correctness/completeness claim (P1). The evaluation supports it only as absence-of-complaint: eight practitioners did not spontaneously report missing factors or misaligned steps, and several said the flow matched how they already think. That is not evidence that analysts reliably reach the same category. The paper explicitly says it did not compare participants' output assignments because interpretative assumptions made outcomes incomparable in the lab setting. Section 6 then concedes that category boundaries can be ambiguous—a deviation might be an exception in one reading and a negative deviation in another—which undercuts the advertised mutual exclusivity. Step 4 also has no aggregation rule for conflicting impact dimensions (compliance vs. performance, etc.); the paper says such trade-offs are typically discussed with management. So the framework underdetermines the final category in exactly the cases where a procedural guide would be most useful.\n\nTo be fair, the authors flag most of this themselves, which is more than many framework papers do. And the practical value can survive the weak P1: as a checklist and orientation device, the framework plausibly reduces ad-hoc variability and helps novices structure conformance-checking follow-up. The expert quotes support that. The citation pattern looks reasonable, with no math or data-fitting to critique.\n\nWho gets value: process-mining researchers working on conformance checking or deviation analysis, and practitioners who want a guided triage of conformance results. It deserves a serious referee—the gap is real and the attempt is honest—but the referee should ask for P1 to be reframed as a design claim, not an established property, and for a follow-up reliability study with measured inter-rater agreement. My verdict: conditionally useful.","headline":"A genuinely new procedural scaffold for deviation triage, but the correctness/completeness claim rests on qualitative self-report rather than demonstrated reliability.","tokens_in":14099,"tokens_out":2800,"would_cite":true,"duration_ms":26623,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that process analysts can replace ad-hoc judgment about whether a deviation is problematic, acceptable, or beneficial with a five-step framework that assigns every deviation to one of seven mutually exclusive…","keywords":["process mining","conformance checking","deviation desirability","process deviation","procedural framework","business process management","action recommendation","qualitative evaluation"],"falsifier":"A controlled application of the framework in a domain not used in its development, such as healthcare or public-sector permitting, where analysts encounter a deviation whose deciding consideration is not one of the eleven factors, or where two analysts applying the same steps to the same traces land in different final categories because of the impact trade-off; finding an unmodelled factor that changes a category, or systematic inter-analyst disagreement, would contradict the completeness claim.","tokens_in":13208,"feed_emoji":"📋","tokens_out":5424,"duration_ms":48510,"temperature":0.7,"pith_summary":"Process mining can find where real executions differ from a prescribed process model, but it cannot say whether those differences are harmless, costly, or beneficial. This paper attempts to close that gap by turning deviation desirability into a repeatable procedure: a five-step framework that guides an analyst through eleven input factors, in a fixed order, to place each deviation into one of seven mutually exclusive categories. The categories are deliberately action-linked, so the output is not just a label but a recommendation to filter out, ignore, prevent, or adopt the behavior. A sympathetic reader would care because the current alternative is ad-hoc expert judgment, which is slow, hard to reproduce across analysts, and difficult to justify. The paper's evaluation with eight practitioners is presented as evidence that the framework achieves its two claimed properties: correct and complete components, and usefulness for orientation.","feed_headline":"Five steps sort process deviations into seven action categories","feed_subtitle":"A new checklist tells process analysts which deviations to ignore, prevent, or adopt instead of judging by gut feel.","key_machinery":"The central object is the procedural framework itself: a five-step decision flow in which eleven input factors act as knockout criteria, meaning a negative answer can assign a final category and stop the flow, and seven output categories are each tied to an action recommendation. The load-bearing design is the order of the steps. First comes data quality; then model correctness and suitability; then case type and control; then compliance, outcome, performance, and standardization; and finally reaction effectiveness and cost. That ordering lets an analyst stop early whenever a deviation is a false alarm or an exception, saving a full impact analysis. The framework's two-part structure separates per-case representational problems from process-wide desirability, and it is this structure that the paper claims is correct, complete, and useful.","core_discovery":"On the paper's own terms, the central discovery is that a deviation's desirability can be assessed procedurally rather than impressionistically. The paper claims that every deviation that conformance checking flags can be routed through two assessment parts: an individual-level part that first excludes false alarms caused by the log, false alarms caused by the model, and justified exceptions, and an aggregated-level part that evaluates true deviations collectively, weighing compliance, outcome, performance, and standardization under both direct and risk-and-opportunity perspectives. If the deviation has positive or negative impact, the analyst then weighs reaction effectiveness against reaction cost, yielding the final categories: false alarm (log), false alarm (model), exception, neutral deviation, positive deviation, negative deviation, or reaction-inefficient deviation. That final assignment is the paper's contribution: a taxonomy plus a procedural path to it.","pith_inferences":["One extension left implicit is that the knockout structure is effectively a decision tree, so the framework could be translated into a software wizard or automated classifier that asks the eleven factors in order; the paper does not specify that translation.","The paper's completeness claim is only as strong as the absence of reported gaps from eight evaluators, so a field study in domains beyond procurement, such as healthcare or public administration, would be the real test of the eleven-factor set.","Because step four leaves the trade-off between compliance and performance to the analyst, the framework likely standardizes the procedure more than the outcome; different organizations could weight those factors differently and arrive at different final categories."],"forward_implications":["If the framework is correct, an analyst can apply the five steps to any conformance deviation and end with one of seven categories, making desirability assessments more replicable across analysts.","The early knockout checks allow many deviations to be set aside as false alarms or exceptions without a full impact analysis, reducing analysis effort in practice.","Each category carries a recommended action, so the assessment output maps directly to decisions: filter out, ignore, prevent, or adopt.","The framework is claimed to be process-agnostic, so a single checklist could be reused across different business domains rather than being rebuilt for each process.","The aggregated-level step brings deviation frequency and collective behavior into desirability assessment, connecting single-case analysis to process-wide improvement efforts."],"supporting_citations":[{"why":"Supplies the process deviation analysis framework and the exception-versus-anomaly distinction that the paper's output taxonomy builds on.","marker":"[12]"},{"why":"Grounds the false-alarm distinction between deviations caused by logging and data quality versus model issues, used in the first two steps.","marker":"[35]"},{"why":"Provides the noise-versus-outlier distinction that informs the log-error check and the data-quality input factor.","marker":"[19]"},{"why":"Contributes the theory of workarounds, which underlies the exception and positive-deviation considerations.","marker":"[4]"},{"why":"Articulates the gap that conformance checking leaves desirability assessment to the analyst, which the framework aims to fill.","marker":"[21]"},{"why":"Supplies the positive-versus-negative deviation distinction that the paper uses in the impact evaluation step.","marker":"[13]"}],"fun_headline_variants":["A two-part procedure classes process deviations into seven types","Systematic framework to rate process deviation desirability","From gut feel to procedure: rating deviation desirability","Five-step guide categorizes process deviation desirability","Seven desirability categories from a step-by-step assessment"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework stands on the assumption that the eleven input factors and their fixed order capture everything an analyst needs to judge desirability, and that analysts can reliably make the qualitative trade-offs the steps demand.","fun_headline_variants_meta":{"raw":{"variants":["A two-part procedure classes process deviations into seven types","Systematic framework to rate process deviation desirability","From gut feel to procedure: rating deviation desirability","Five-step guide categorizes process deviation desirability","Seven desirability categories from a step-by-step assessment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000538,"raw_usage":{"total_tokens":2557,"prompt_tokens":893,"completion_tokens":1664,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":1586}},"tokens_in":509,"tokens_out":1664,"duration_ms":15324,"temperature":1.0,"reasoning_tokens":1586,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:03:20.806472+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled application of the framework in a domain not used in its development, such as healthcare or public-sector permitting, where analysts encounter a deviation whose deciding consideration is not one of the eleven factors, or where two analysts applying the same steps to the same traces land in different final categories because of the impact trade-off; finding an unmodelled factor that changes a category, or systematic inter-analyst disagreement, would contradict the completeness claim.","supporting_citations":[{"cited_title":"In: BPM Workshops","cited_arxiv_id":null,"evidence_quote":"Supplies the process deviation analysis framework and the exception-versus-anomaly distinction that the paper's output taxonomy builds on."},{"cited_title":"J Biomed Inform 85, 155–167 (2018)","cited_arxiv_id":null,"evidence_quote":"Grounds the false-alarm distinction between deviations caused by logging and data quality versus model issues, used in the first two steps."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the noise-versus-outlier distinction that informs the log-error check and the data-quality input factor."},{"cited_title":"CAIS 34, 1041–1066 (2014)","cited_arxiv_id":null,"evidence_quote":"Contributes the theory of workarounds, which underlies the exception and positive-deviation considerations."},{"cited_title":"In: ICPM","cited_arxiv_id":null,"evidence_quote":"Articulates the gap that conformance checking leaves desirability assessment to the analyst, which the framework aims to fill."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the positive-versus-negative deviation distinction that the paper uses in the impact evaluation step."}],"review_version":1}