{"id":"d96c5a44-a88b-4615-ad0a-aba39e02328e","arxiv_id":"2606.24619","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"User study finds that natural language ambiguities in competency questions hinder reliable ontology verification even with LLM assistance, recommending pre-publication refinement tools.","lead":"The paper reports an experiment with 19 participants using an LLM assistant (OE-Assist) to perform competency question verification on 20 ontology tasks. It concludes that ambiguities and complexity in CQs require dedicated refinement tools before publication.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Single-condition study with N=19 cannot isolate whether refinement tool is necessary vs. other factors","rationale":"The reader's weakest_assumption directly identifies the representativeness/generalization risk; the additional design gap (no control for the proposed intervention) is the precise reason the necessity conclusion remains under-supported even if the observed challenges are real.","tokens_in":1630,"tokens_out":318,"duration_ms":10022,"concrete_test":"Re-run the 20 tasks with the same 19 participants after an independent expert panel produces refined versions of the original CQs; compare per-task error rate, time, and inter-rater consistency between original and refined conditions under identical OE-Assist usage. If the refined condition shows no statistically significant improvement (p>0.05 or effect size <0.3), the necessity claim does not follow from the experiment.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the 19-participant / 20-task experiment with OE-Assist demonstrates the necessity of a pre-publication CQ refinement tool. For this to hold, the observed ambiguities and errors must be shown to stem specifically from unrefined CQs (rather than from OE-Assist limitations, task selection, or participant variability) and to be reliably reduced by refinement. The reported design appears to be a single arm with no baseline (no-assistant or pre-refined CQs), no quantitative metrics of verification correctness, and no statistical test of generalizability, leaving the necessity inference unsupported by the data.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript reports an empirical user study in which 19 participants performed CQ-verification on 20 tasks using the OE-Assist LLM assistant. It identifies challenges arising from ambiguities and excessive complexity in natural-language CQs and concludes that these results demonstrate the necessity of a dedicated pre-publication CQ refinement tool to improve subsequent ontology engineering phases.","tokens_in":1762,"tokens_out":369,"duration_ms":26439,"significance":"If the central claim were supported by appropriate controls and metrics, the work would provide a useful empirical illustration of practical difficulties in CQ-based ontology verification and could motivate tool development in the ontology engineering community. The study design incorporates real participants interacting with an LLM assistant, which supplies a concrete, practice-oriented data point.","major_comments":[{"comment":"Abstract: the single-condition design (19 participants, 20 tasks, OE-Assist only) supplies no baseline arm (e.g., pre-refined CQs or no assistant) and reports no quantitative correctness metrics, statistical tests, or raw data. Consequently the observed errors cannot be attributed specifically to unrefined CQs, which is load-bearing for the necessity claim.","section":"Abstract"},{"comment":"Experimental setup / Results: without a control condition or within-subject comparison, the data cannot test whether the reported ambiguities and complexity would be reliably reduced by a refinement tool, leaving the causal inference that such a tool is necessary unsupported.","section":"Experimental setup / Results"}],"minor_comments":[{"comment":"Abstract: participant demographics, task selection criteria, and exact performance measures are omitted, reducing the reader's ability to assess representativeness.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"Thank you for the referee's insightful comments. We recognize the limitations of our single-condition study design and will revise the manuscript to clarify the exploratory nature of the work, qualify our conclusions regarding the necessity of a CQ refinement tool, and add a limitations section to address the lack of baseline comparisons and quantitative metrics.","responses":[{"response":"The referee correctly identifies that our study employs a single-condition design without a baseline. This was intentional as the goal was to investigate challenges in CQ verification as it is currently practiced with LLM assistance, rather than to evaluate the impact of refinement. The analysis was qualitative, focusing on participant feedback and observed issues, which explains the absence of quantitative metrics and statistical tests. We will revise the abstract to better describe the study as exploratory and to moderate the claim about demonstrating necessity, instead highlighting the observed challenges as motivation for tool development. Raw data can be made available upon request in a revision if it aids transparency.","revision_made":"partial","referee_comment":"[Abstract] Abstract: the single-condition design (19 participants, 20 tasks, OE-Assist only) supplies no baseline arm (e.g., pre-refined CQs or no assistant) and reports no quantitative correctness metrics, statistical tests, or raw data. Consequently the observed errors cannot be attributed specifically to unrefined CQs, which is load-bearing for the necessity claim."},{"response":"We concur that the current data does not support a causal claim about the effectiveness of a refinement tool, as no comparison was made. The recommendation for such a tool stems from the identification of specific ambiguities and complexities that hinder verification, suggesting that preemptive refinement could mitigate these. In the revised manuscript, we will reframe the conclusions to present this as a motivated recommendation for future research and tool building, rather than an empirically proven necessity. We will explicitly state in a new limitations paragraph that controlled experiments are needed to confirm the benefits of CQ refinement tools.","revision_made":"partial","referee_comment":"[Experimental setup / Results] Experimental setup / Results: without a control condition or within-subject comparison, the data cannot test whether the reported ambiguities and complexity would be reliably reduced by a refinement tool, leaving the causal inference that such a tool is necessary unsupported."}],"tokens_in":1247,"tokens_out":486,"duration_ms":25585,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper runs a 19-person study where participants used an LLM assistant called OE-Assist to verify competency questions against ontologies. It reports that people ran into ambiguities and complexity, then concludes a pre-publication refinement tool is needed.\n\nWhat the work actually does is document some practical friction points that ontology engineers already know exist when CQs are poorly worded. The participant tasks and the assistant are described at a high level, and the authors correctly flag that linguistic issues affect later modeling steps.\n\nThe soft spot is the leap from observations to necessity. There is no control condition, no pre/post refinement comparison, no correctness scores, and no statistical checks. With N=19 and 20 tasks in one arm only, you cannot separate problems caused by raw CQs from problems caused by the assistant itself, task choice, or individual differences. The abstract supplies none of the numbers or demographics that would let a reader judge whether the pattern is reliable.\n\nThis is the kind of incremental tooling note that might interest a narrow group working on ontology evaluation interfaces. It does not yet give enough evidence to change practice or to justify a dedicated refinement component.\n\nI would not bring it to reading group and would not cite it. A serious editor should desk-reject rather than send it out, because the central inference rests on an unsupported design.","headline":"Small single-arm study with no metrics or baseline can't support the claim that a CQ refinement tool is necessary.","tokens_in":2221,"tokens_out":336,"would_cite":false,"duration_ms":10999,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Competency questions with ambiguities and excessive complexity hinder reliable ontology verification and require a dedicated refinement tool before publication.","keywords":["Competency Questions","Ontology Verification","CQ Verification","LLM Assistant","Ontology Engineering","Ambiguity","Complexity","Refinement Tool"],"falsifier":"A controlled comparison in which the same 20 tasks are rerun after systematic refinement of the CQs, measuring whether the rate of inconsistent modelling decisions drops.","tokens_in":2563,"feed_emoji":"⚠️","tokens_out":580,"duration_ms":21157,"temperature":0.7,"pith_summary":"The paper examines what makes competency questions challenging during the verification process that checks whether an ontology meets its intended purpose. Experiments involved 19 participants completing 20 tasks with an LLM-based assistant, revealing that linguistic nuances often produce inconsistent modelling decisions. The results establish that these issues persist even with assistance and point to the need for a tool that refines CQs prior to their use in evaluation. A sympathetic reader would see this as a practical step to reduce errors in later ontology engineering phases.","feed_headline":"Ambiguous CQs complicate ontology verification","feed_subtitle":"Experiments with an LLM assistant show that linguistic issues lead to inconsistent results, indicating a need for pre-publication refinement","key_machinery":"OE-Assist, the LLM assistant deployed to support participants in CQ-verification tasks, used to surface specific interpretation challenges across the 20 tasks.","core_discovery":"CQ-verification is time-consuming and error-prone because it requires careful interpretation of linguistic nuances and precise alignment with formal ontology constructs; ambiguities and complexity in CQs lead to inconsistent modelling decisions, and the experiments demonstrate the necessity of a tool to refine CQs before publishing them to avoid these problems in the ontology engineering process.","pith_inferences":["A similar refinement step could apply to other natural-language specifications used in knowledge engineering.","Detection rules for common CQ ambiguities might be added to existing ontology tools to automate part of the process.","Refinement could be tested as a standard checkpoint before any CQ-based evaluation begins."],"forward_implications":["Refined CQs produce more consistent alignment between natural language questions and formal ontology constructs.","Unrevised CQs increase the likelihood of error-prone and time-consuming verification outcomes.","A pre-publication refinement step reduces ambiguity that otherwise propagates into later ontology engineering phases.","LLM assistance alone does not eliminate the need for prior CQ refinement."],"fun_headline_variants":["Linguistic issues hinder CQ verification","Ambiguities in CQs lead to inconsistent results","Refine CQs to avoid verification problems","Experiments confirm need for CQ refinement"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The challenges observed with 19 participants across 20 tasks using OE-Assist are representative of typical CQ verification difficulties and would persist without a dedicated refinement tool.","fun_headline_variants_meta":{"raw":{"variants":["Linguistic issues hinder CQ verification","Ambiguities in CQs lead to inconsistent results","Refine CQs to avoid verification problems","Experiments confirm need for CQ refinement"]},"model":"grok-4.3","cost_usd":0.00565,"raw_usage":{"total_tokens":2663,"prompt_tokens":591,"num_sources_used":0,"completion_tokens":51,"cost_in_usd_ticks":56499500,"prompt_tokens_details":{"text_tokens":591,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2021,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":591,"tokens_out":51,"duration_ms":13982,"temperature":1.0,"reasoning_tokens":2021,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-25T23:35:35.513814+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled comparison in which the same 20 tasks are rerun after systematic refinement of the CQs, measuring whether the rate of inconsistent modelling decisions drops.","supporting_citations":[],"review_version":1}