{"id":"9367b4f2-78ca-447b-9729-9c51d1ecd17f","arxiv_id":"2504.17295","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a real insurance claims pipeline, an LLM identified claim parts in 27.6% of claims versus 1.8% for human handlers, scaling that step 14-fold while shifting the bottleneck to the investigation stage.","lead":"A Swedish insurance company deployed GPT-4o to spot claim parts that need special investigation, replacing a manual step that had become a bottleneck. Process mining on five months of production data shows the AI found many more claim parts than humans did, but the extra workload then clogged the downstream investigation stage.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 1420% scaling claim equates 1034 raw AI predictions with 68 human reports, but production precision is never measured; only 23 of 1034 predictions led to validated investigations, and the 1.82% vs 27.62% gap is inconsistent with the 70% human recall benchmark.","rationale":"The reader's weakest assumption concerns OCEL 2.0 keeping only the most recent claim note before scan. That is a real limitation, but it is not the most load-bearing issue: the 1034 pCP count is a production event count and does not depend on which notes were retained in the OCEL. The scaling numerator itself is unvalidated. My concern is different: raw LLM positive outputs are compared with expert human reports as if they were the same kind of count. The paper provides a balanced-dataset precision estimate and 26 investigated-case outcomes, but no production ground truth for the 1034 predictions. The internal inconsistency between 1.82% human production yield and the 70% human recall benchmark makes this especially acute. I keep the reader's CONDITIONAL verdict (no change), but the condition should be a production precision/recall audit, not log-extraction fidelity. I credit the paper for deploying in production, transparently describing extraction limits, and validating the 26 cases with stakeholders; the concern is about the quantitative scaling number, not the qualitative finding that AI changes process dynamics.","tokens_in":11214,"tokens_out":9110,"duration_ms":93911,"concrete_test":"Select a random sample of 200 of the 1034 claims with AI-predicted claim parts and 100 claims with no AI prediction from the production log; have claim part investigators, blinded to AI output, label each claim for presence of the relevant claim parts using the same criteria as the 26 investigation cases. Compute production precision and recall for AI, and apply the same gold standard to a sample of the 68 human-reported claims. If precision times 1034 remains more than 14 times the human-validated count, the scaling headline holds; otherwise report a confidence interval for the true AI-identified count and revise the 1420% claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 Q2 (Fig. 6) reports 1034 AI-identified claim parts and computes a 1420% scaling versus 68 human reports. This comparison is the load-bearing premise of the paper. The pCP events are raw LLM outputs; the rCP events are expert reports. The model's precision (0.81 English, 0.83 Finnish for v5) was measured on a balanced, manually labeled dataset (Section 3.2, Table 1), not on production data. There is no sample of the 1034 predictions checked by investigators; the only production-level confirmation is that AI identified 23 of the 26 investigated cases, which is a recall-like figure for investigated cases, not a precision estimate for the 1034 predictions. The paper's own numbers expose an internal inconsistency: if the AI's 27.62% prevalence is correct, the humans' 1.82% corresponds to roughly 6.6% production recall, far below the 70% human recall baseline stated in Section 3.2. Either the 68 count is not a recall-limited measure of human identification, or many of the 1034 predictions are false positives. The phrase 'confirmed by the business' is informal and does not quantify precision. The exact 1420% figure requires near-perfect production precision, which is not demonstrated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a real-world case study at If P&C Insurance in which a GPT-4o-based LLM was deployed in production to identify insurance claim parts, replacing a manual identification process that had been identified as a scalability bottleneck. The authors follow a combined BPR, CRISP-DM, and OCPM2 methodology, extract an OCEL 2.0 event log over five months covering 3,743 claims, and use object-centric and flat process mining to compare human and AI identification behavior. Their central empirical claim is that the AI scaled claim-part identification by roughly 1420%: humans reported 68 claim parts (1.82% of claims) while the AI predicted 1,034 claim parts (27.62% of claims). They further report that downstream claim-part investigation became the new bottleneck, with only 26 investigation cases created, and that among those 26 cases, AI alone identified 5, humans alone identified 3, and both identified 18. The paper also discusses lessons learned about process-mining visualization and stakeholder communication, and concludes that AI-driven automation can shift bottlenecks rather than eliminate them.","tokens_in":11428,"tokens_out":4269,"duration_ms":42307,"significance":"If the empirical counts are reliable, the paper provides a valuable, rare, real-world demonstration of using object-centric process mining to quantify the side effects of introducing an LLM into a knowledge-intensive insurance process. The observation that AI can remove one bottleneck only to create another, and that OCPM can make this trade-off visible from production event logs, is practically important and would be of interest to the BPM and process-mining communities. The paper is transparent about several limitations, including the OCEL 2.0 restriction on representing expired relations and the difficulty of communicating object-centric models to stakeholders. The main strength is the concrete, company-scale empirical setting and the honest reporting of downstream constraints. However, the central quantitative claim about a 1420% scaling depends on treating raw LLM predictions as equivalent to human-identified claim parts without a production precision measurement, which limits the force of the conclusions as currently stated.","major_comments":[{"comment":"The 1420% scaling figure compares 1,034 pCP events (raw LLM predictions) with 68 rCP events (hand-reported claim parts). The pCP events are model outputs, not validated claim parts, and the paper provides no production precision estimate for the 1,034 predictions. The only production-level validation reported is that AI was involved in 23 of the 26 investigation cases, which is a recall-like figure for the investigated subset, not a precision measure for all predictions. Moreover, the numbers are internally tension-prone: if the stated 70% human recall baseline applies in production, then 68 human reports imply roughly 97 true claim parts, so the AI's 1,034 predictions would imply a very high false-positive rate; conversely, if 1,034 is the correct count of true claim parts, then human production recall would be 68/1034 ≈ 6.6%, far below the 70% baseline. The phrase \"confirmed by the business\" does not quantify precision. The authors should either reframe the claim as \"candidate lead generation\" or provide a production precision evaluation (e.g., a random sample of the 1,034 predictions reviewed by investigators) before asserting that the process has scaled by 1420%.","section":"Section 3.3, Q2 and Fig. 6; Section 3.2, Table 1"},{"comment":"The extracted OCEL 2.0 log retains only the most recent claim note prior to the scan activity because OCEL 2.0 cannot represent expired object-to-object relations. Since the AI in production reads claim descriptions and notes, and the log is the basis for the pCP counts and the Q3/Q4 complementarity results, it is unclear whether the documented 1,034 predictions and the 3/5/18 breakdown reflect the data the AI actually used or an artifact of this one-note snapshot. The authors acknowledge the limitation, but they should clarify whether the pCP event counts come directly from the production AI system (which would be unaffected by the OCEL note-filtering) or from the reconstructed OCEL, and they should discuss how sensitive the reported ratios are to this filtering choice.","section":"Section 3.3, Implementation and Log Extraction"},{"comment":"The complementarity findings (AI missed 3, humans missed 5, both identified 18) are computed only over the 26 investigation cases, which is a small, deliberately selected subset of the 3,743 claims (investigators choose cases based on business criteria). As reported, the numbers describe the investigated set, not the full population of claims or all identified claim parts. The paper should state this limitation directly in the Q3/Q4 analysis and temper any conclusion about AI's general added value in identification, since the 26 cases are not a random or representative sample of the claims processed during the study period.","section":"Section 3.3, Q3/Q4 and Fig. 9"}],"minor_comments":[{"comment":"The figure shows pCP with frequency 1,069 and a loop, while the text correctly notes that the incoming flow shows 1,034 unique predictions; adding an explicit label for the unique prediction count on the figure would reduce ambiguity.","section":"Section 3.3, Fig. 4"},{"comment":"Reference [14] contains the typo \"Accpeted\" and reference [17] says \"Accepted in BPMDS 2025\" without a year; these should be corrected for consistency.","section":"References"},{"comment":"The sentence \"This limitation does not exist in all OCED formats\" appears to contain a typo; it should presumably be \"OCEL\" or \"object-centric event data formats.\"","section":"Section 3.3, Implementation and Log Extraction"},{"comment":"The company name is rendered with a typographical oddity: \"If P&C Insurance , a property...\" has a stray space before the comma; this should be cleaned up.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a legitimate experience report and the OCPM-based bottleneck-shift observation is a useful contribution. The main concern is internal consistency between the reported human recall baseline, the 68 human reports, and the 1,034 AI predictions; the authors need to address this head-on, either by adding a production precision measurement or by demoting the 1420% claim to a candidate-generation throughput claim. A substantial revision that reframes the central claim and adds the necessary caveats would make the paper publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core news: this is a genuine production deployment of an LLM for insurance claim part identification, with five months of event log data covering 3743 claims while humans and AI ran in parallel. The paper applies OCPM to compare the coexisting variants, and the qualitative finding is solid: AI removed the identification bottleneck and pushed pressure downstream to the investigation step, where only 26 investigation cases were created despite 1034 AI predictions. That shift is a useful, believable lesson for practitioners. The paper is also transparent about the OCEL note-filtering limitation and the message queue duplication in the log, and it credits the business experts who validated the process-level findings. The use of OCPM to attribute investigations to the correct variant (AI, human, or both) is a genuine extension, not just a toy example.\n\nWhere the paper is soft is exactly where the stress-test note lands. The 1420% scaling figure divides 1034 pCP events by 68 rCP events, but those are not commensurate quantities. pCP is a raw LLM output; rCP is a human formal report. The paper does not sample the 1034 predictions to estimate production precision. The only production-level validation is that AI predicted 23 of the 26 investigated cases, which is a recall-like figure for a very small and possibly cherry-picked set, not a precision estimate. The paper also creates an internal inconsistency: if the 27.62% claim prevalence is correct, the human production rate of 1.82% implies a far lower recall than the 70% benchmark cited for the evaluation set. One can reconcile that by noting the benchmark was on a balanced labeled dataset while production involves busy claim handlers with different incentives, but the paper does not do that reconciliation, and the 1420% number is therefore not a clean measurement of process scaling. This is the load-bearing weakness, and it is real.\n\nOther concerns are minor: the 3/5/18 complementarity split rests on 26 investigation cases, so it is anecdotal; the log retains only the most recent note per claim, which the authors acknowledge; no anonymized log or code is provided. The citation pattern is fine; the self-citations to OCPM2 and tEKG define the toolkit and do not determine the empirical result.\n\nWho should read this: process mining practitioners and researchers interested in AI-in-the-loop field studies, and BPM people thinking about automation bottlenecks. It deserves a serious referee, but the referee should push hard on the scaling claim. My recommendation: engage with the paper, and if you cite it, cite the bottleneck-shift lesson, not the 1420% number.","headline":"A real production case study worth reading, but the headline 1420% scaling figure is not demonstrated because it compares raw AI predictions to human reports without measuring AI precision.","tokens_in":12020,"tokens_out":2636,"would_cite":true,"duration_ms":28649,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An LLM deployed in production at an insurer identified 1,034 claim parts over five months versus 68 identified by human handlers, a 1,420% increase, while the investigation stage became the new bottleneck.","keywords":["AI-Driven Automation","Business Process Reengineering","Digital Transformation","Business Process Management","Object-Centric Process Mining","Large Language Models","Claims Processing","Scalability"],"falsifier":"Re-extract the same 3,743 claims with full claim-note history, or with alternative snapshots of the notes available at scan time, and recompute the AI and human identification counts and the attribution of the 26 investigation cases; if those numbers move materially, the 1,420% scaling figure and the 3/18/5 complementarity are not robust.","tokens_in":10977,"feed_emoji":"🤖","tokens_out":5941,"duration_ms":50836,"temperature":0.7,"pith_summary":"Faced with rising claim volumes, an insurance company replaced a manual, knowledge-intensive screening step with a large language model that scans claim descriptions and notes and predicts which claims contain claim parts requiring special handling. The paper uses object-centric process mining on a production event log covering 3,743 claims over five months to compare the old and new process variants running in parallel. It reports that the AI identified 1,034 claim parts while human handlers identified 68, a 1,420% scale-up, and that the ratio of claims with claim parts rose from 1.82% to 27.62%, in line with what business experts expected. The same analysis shows the next step did not scale: only 26 investigation cases were created, because the number of investigators had been sized for the lower volume. The paper's central point is that AI can remove one bottleneck only to create another, and that object-centric logs make that bottleneck shift visible and attributable to each process variant.","feed_headline":"AI scaled claim-part detection by 1,420%","feed_subtitle":"A production case study at an insurer shows the new bottleneck moved downstream: only 26 investigations were opened.","key_machinery":"The central object is the Object-Centric Event Log (OCEL 2.0), a data format in which each event can relate to several object types instead of a single case; here the object types are Customer, Claim, Claim Note, Claim Part, AI Model, Claim Handler, and Claim Part Investigator. Two AI-specific activities, scan claim and predict claim part, record when the model reads a claim's description and notes and when it flags a claim part. Analysis proceeds through Object-Centric Directly-Follows Graphs (OC-DFGs) and Object-Centric Petri Nets, with drill-down, unfolding, and flattening operations used to separate human-only from AI-involved paths. The mechanism that carries the argument is the ability to filter and unfold the log by role and by AI involvement, which lets the authors attribute each investigation case to humans, the AI, or both; the paper also notes a limitation, that OCEL 2.0 cannot represent expired object-to-object relations, so the log keeps only the most recent claim note before each scan.","core_discovery":"On the paper's own terms, the discovery is that a production large language model can scale claim-part identification far beyond human throughput while maintaining acceptable quality, and that object-centric process mining can quantify the consequences. The model identified claim parts for 1,034 of 3,743 claims (27.62%), versus 68 claims (1.82%) identified by human handlers, whom the company's baseline analysis estimated miss 30% of true claim parts. The paper reports a 1,420% scaling of the identification step and notes that among the 26 claim parts opened for investigation, 18 were found by both AI and humans, 5 were found only by AI, and 3 were found only by humans. It then argues that this success exposed a new constraint: claim part investigators opened only 26 cases, so the downstream investigation stage, not identification, now limits the claims management process. The paper concludes that scaling a single step does not automatically add business value and that end-to-end process redesign is needed when AI is introduced.","pith_inferences":["The same pattern likely generalizes to other AI-screening applications, such as medical triage, fraud alerting, or content moderation: a high-recall model increases queue volume at the human review stage, and the new queue location is visible only if the log preserves the object relations that connect screening outputs to downstream work.","A direct test of the paper's interpretation would be to measure investigation cycle time and investigator utilization before and after deployment; a growing queue with rising wait times would confirm the bottleneck shift, while a flat queue would suggest the 26-investigation count reflects business selection criteria rather than capacity constraints.","Because the extracted log keeps only one claim note per claim before the AI scan, re-running the same comparison on a temporal event knowledge graph that preserves note history would show whether the 1,034-versus-68 contrast and the 3/18/5 complementarity counts depend on the snapshot assumption."],"forward_implications":["Automating a single bottleneck step in a business process shifts rather than removes the constraint; downstream capacity must be planned in the same reengineering effort.","Object-centric event logs make it possible to evaluate AI and human process variants during a gradual transition, because outcomes can be attributed to each variant by filtering on activity and object type.","Flattened, single-case graphs are more useful than full object-centric models for communicating results to business stakeholders, even when the object-centric representation is what enables the analysis.","Recall, rather than F1-score, is the right quality target when missed cases are far more costly than false positives, as in claim-part identification.","If the AI's identification capacity is used fully, the organization must hire or train more investigators, automate parts of the investigation step, or redesign the workflow, or the added detections will not turn into handled cases."],"supporting_citations":[{"why":"Defines object-centric process mining and the Object-Centric Directly-Follows Graph used to analyze AI and human variants.","marker":"[22]"},{"why":"Provides the OCEL 2.0 specification whose relation model shapes the log extraction and whose lack of expired-relation support motivates the one-note snapshot.","marker":"[3]"},{"why":"Supplies the OCPM2 methodology the evaluation follows.","marker":"[17]"},{"why":"Provides the business process reengineering framework that structures the case study's phases.","marker":"[18]"},{"why":"Supplies the CRISP-DM process used to develop and validate the LLM model.","marker":"[7]"},{"why":"Establishes the few-shot learning capability of large language models that the claim-part identification task relies on.","marker":"[5]"},{"why":"The PM2 process mining methodology that OCPM2 extends.","marker":"[24]"},{"why":"Documents the structured output feature that forces the model's predictions into a consistent schema, making the predict claim part events reliable.","marker":"[20]"}],"fun_headline_variants":["LLM scales claim ID 1,420%, but bottleneck shifts downstream","AI finds 1,420% more claim parts; investigations now the limit","Insurance case: AI scales identification, exposes new bottleneck","OCPM shows AI's 1,420% gain; downstream still caps process","AI in insurance: 1,420% more identifications, only 26 investigations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the single most recent claim note retained in the log before each AI scan faithfully represents the claim text the AI actually read in production, so that the counts of AI-identified and human-identified claim parts are accurate rather than artifacts of log extraction.","fun_headline_variants_meta":{"raw":{"variants":["LLM scales claim ID 1,420%, but bottleneck shifts downstream","AI finds 1,420% more claim parts; investigations now the limit","Insurance case: AI scales identification, exposes new bottleneck","OCPM shows AI's 1,420% gain; downstream still caps process","AI in insurance: 1,420% more identifications, only 26 investigations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000348,"raw_usage":{"total_tokens":1906,"prompt_tokens":948,"completion_tokens":958,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":860}},"tokens_in":564,"tokens_out":958,"duration_ms":7339,"temperature":1.0,"reasoning_tokens":860,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:43:22.012047+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-extract the same 3,743 claims with full claim-note history, or with alternative snapshots of the notes available at scan time, and recompute the AI and human identification counts and the attribution of the 26 investigation cases; if those numbers move materially, the 1,420% scaling figure and the 3/18/5 complementarity are not robust.","supporting_citations":[{"cited_title":"Object-centric process mining: dealing with divergence and con- vergence in event data","cited_arxiv_id":null,"evidence_quote":"Defines object-centric process mining and the Object-Centric Directly-Follows Graph used to analyze AI and human variants."},{"cited_title":"OCPM 2: Ex- tending the Process Mining Methodology for Object-Centric Event Data Extraction, 2025","cited_arxiv_id":null,"evidence_quote":"Supplies the OCPM2 methodology the evaluation follows."},{"cited_title":"Business process reengineering: A theoretical framework and an integrated model","cited_arxiv_id":null,"evidence_quote":"Provides the business process reengineering framework that structures the case study's phases."},{"cited_title":"Crisp-dm 1.0: Step-by-step data mining guide","cited_arxiv_id":null,"evidence_quote":"Supplies the CRISP-DM process used to develop and validate the LLM model."},{"cited_title":"van Eck, Xixi Lu, Sander J","cited_arxiv_id":null,"evidence_quote":"The PM2 process mining methodology that OCPM2 extends."},{"cited_title":"Structured outputs - openai api","cited_arxiv_id":null,"evidence_quote":"Documents the structured output feature that forces the model's predictions into a consistent schema, making the predict claim part events reliable."}],"review_version":1}