{"id":"6e9c6bcd-4cc4-4043-87d2-7d5fe950b51d","arxiv_id":"2507.10228","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper outlines, but does not implement or validate, a framework that merges the artefact-based method AMDiRE with the perspective-based method PerSpecML to specify trustworthy AI requirements.","lead":"This paper proposes combining two existing requirements engineering methods, AMDiRE and PerSpecML, into one framework that turns high-level trustworthy AI goals into structured, traceable specifications. It is a vision paper: the combination is described and illustrated with a fictional hiring platform example, but it is not yet built or tested.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central bridging claim rests on an unexamined assumption that PerSpecML concerns can be systematically mapped onto AMDiRE artefacts; the paper's only illustration is hand-crafted and contains a malformed requirement.","rationale":"The paper is explicitly a vision paper. Its abstract and Section III frame the contribution as a proposal ('we envision', 'we plan to'), and the conclusion states that development 'can only yield meaningful results if continuously evaluated based in practical settings.' There is no implemented artifact, dataset, or falsifiable outcome, so UNVERDICTED is the right verdict. The source methods themselves are established and tool-supported, so the risk is not in AMDiRE or PerSpecML individually but in the proposed integration. The single load-bearing technical assumption is that a systematic mapping exists between PerSpecML's concern task model and AMDiRE's artefact content structures at compatible granularity. Section III-B lists this as a planned step, and Table II is the only illustration; it is hand-crafted and includes a malformed sample requirement, so it cannot support more than a plausibility claim. My concrete test would force the mapping to be stated precisely for one regulatory obligation and would reveal whether the integration is a straightforward extension or requires a new intermediate model. This does not change the reader's verdict: the concern is about future feasibility, not about a false claim in the paper.","tokens_in":6005,"tokens_out":3262,"duration_ms":40405,"concrete_test":"Select a single EU AI Act obligation, e.g. Article 14 human oversight, and the PerSpecML concerns it triggers. Populate every content element of the AMDiRE Context Specification and Requirements Specification that the framework would require, using only the mappings implied by Table II. If any concern cannot be expressed without weakening the artefact's content structure, or if the resulting text contains a requirement that is syntactically malformed or untraceable to the regulation, the systematic bridging claim fails. This check is feasible on paper before any tool is built.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is a feasibility vision: extending PerSpecML's concern catalogue with regulatory requirements and mapping concerns onto AMDiRE's artefact model yields a concern-aware, artefact-driven pathway for trustworthy AI requirements. The load-bearing assumption is that this mapping is systematic and preserves granularity. Section III-B lists 'Bridging to artefact models' as a future step but provides no metamodel-level compatibility analysis, no mapping rules, and no constraints on granularity. The only evidence offered, Table II, is hand-crafted. It contains one entry that is not a well-formed requirement: 'The system architecture shall performance under input perturbations' (missing verb). The table maps 'Model Robustness' to 'System Specification' without explaining which content elements of that AMDiRE artefact receive which information, or how non-functional, emergent concerns such as fairness or human oversight translate into the artefact's predefined content structure. If concerns and artefact content elements are incommensurable at the required granularity, the integration would require inventing a new intermediate model, which is not described. Since the paper explicitly confines itself to a vision and its conclusion states that meaningful results require practical evaluation, this concern does not falsify the paper; it identifies the key condition that any future implementation must satisfy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This short vision paper proposes integrating two existing requirements engineering approaches: AMDiRE, an artefact-based RE method, and PerSpecML, a perspective-based method for ML-enabled systems. The motivation is that AMDiRE offers structured, traceable artefacts but was designed for deterministic systems, while PerSpecML provides stakeholder-driven concern elicitation for ML contexts including trustworthiness-related concerns. The authors envision a framework in which PerSpecML's concern catalog is extended with regulatory requirements (e.g., the EU AI Act), mapped onto AMDiRE's artefact model, and supported by a tool. The paper illustrates the idea with a hiring-platform example (Tables I and II), reviews related work, and concludes that practical evaluation in industry collaborations is required for the vision to produce meaningful results.","tokens_in":6147,"tokens_out":5283,"duration_ms":62605,"significance":"If the envisioned integration is realized, it would address a genuine gap: translating high-level trustworthiness principles into concrete, traceable RE artefacts. The paper builds transparently on the authors' prior methods (AMDiRE and PerSpecML), which is appropriate for a vision paper, and it explicitly acknowledges that the mappings require validation. The clearest strength is the clear articulation of complementary strengths and the structured set of research directions. The chief weakness is that the only concrete illustration of the central bridging mechanism is a hand-crafted table with a malformed entry, so the reader cannot yet assess whether the mapping is feasible at the required granularity.","major_comments":[{"comment":"The only concrete illustration of the concern-to-artefact mapping contains a malformed requirement entry: 'The system architecture shall performance under input perturbations' is missing the main verb and is not a well-formed requirement. This is a load-bearing defect because Table II is the paper's sole evidence that the envisioned operationalization can work. The authors must correct the example and ideally explain which specific content elements of the target artefact (e.g., the quality model of the System Specification) receive the concern-related information. Without a coherent example, the feasibility of the bridging step remains unsubstantiated.","section":"Section III-C, Table II"},{"comment":"The paper asserts that the extended concern model will be mapped to AMDiRE's artefacts, but it provides no method or criteria for this mapping and no analysis of whether the two metamodels are compatible at the required level of granularity. Non-functional, emergent concerns such as fairness or human oversight may not correspond naturally to the predefined content structures of AMDiRE's context, requirements, or system specification artefacts. Since this compatibility is the central enabling assumption of the proposed framework, the paper should either provide a proof-of-concept mapping for at least one concern-artefact pair (beyond the flawed example) or explicitly position the mapping as an open research question rather than as a planned step. As written, the claim that the approaches 'can and should' be integrated overstates what is currently shown.","section":"Section III-B, 'Bridging to artefact models'"}],"minor_comments":[{"comment":"The phrase 'let along' should be 'let alone'.","section":"Section I"},{"comment":"The sentence 'Governmental bodies and institutions at, such as the ones in the European Union' contains a misplaced 'at'; it should read 'institutions, such as those in the European Union'.","section":"Section II-A"},{"comment":"The sentence 'since of them (e.g., decision trees, linear regression) are inherently more explainable' should read 'since some of them'.","section":"Section II-C"},{"comment":"The first column of Table II is labeled 'Concern', but the entry 'Stakeholder Roles' is not a trustworthiness concern; it is a modeling element, which makes the mapping unclear. Please either remove it or rename the column to something broader such as 'Element'.","section":"Section III-C, Table II"},{"comment":"There are two language issues in the concluding paragraph: 'an unique role' should be 'a unique role', and 'will hopefully contributing' should be 'will hopefully contribute'.","section":"Section V"}],"recommendation":"major_revision","confidential_remarks":"The paper is more appropriate for a vision/position track than a full research paper, but even for a vision paper the illustrative example needs to be technically correct. The self-citation is transparent and not a concern. The core issue is that the central bridging step is asserted rather than demonstrated; the authors should be encouraged to add a small, correct worked example or an initial metamodel compatibility analysis. This is fixable within the scope of a revised paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a short vision paper, and it is honestly labeled as one. The useful idea is pairing AMDiRE's artefact-centric structure with PerSpecML's concern catalog to give requirements engineers a traceable route from trustworthiness principles to concrete specifications. That is a plausible and bounded contribution to the RE subfield, and the paper does not overclaim: it explicitly says validation is future work and calls for practical evaluation. What the paper does well is positioning. The background on trustworthy AI is compact, the related work is relevant, and the authors correctly identify that neither source method alone covers the gap. The illustrative hiring-platform example, despite being hand-crafted, at least makes the vision concrete. The self-citation is transparent and not by itself a problem, since both methods are published and the authors know them deeply. The soft spots are real but not disqualifying. The load-bearing assumption is that PerSpecML's concerns can be systematically mapped onto AMDiRE's artefact content structures, preserving granularity and covering non-functional concerns like fairness or human oversight. The paper does not analyze this compatibility; it only asserts the mapping in Table II. That table also contains a genuinely sloppy entry: 'The system architecture shall performance under input perturbations' is not a well-formed requirement. That is a credibility ding. But the paper itself lists 'bridging to artefact models' as a future step and concludes that meaningful results require empirical evaluation, so the stress-test's concern identifies the key condition rather than a contradiction internal to the paper. Who is this for? Researchers already working with AMDiRE or PerSpecML, and RE folks thinking about how to operationalize the EU AI Act. It is workshop-level material, not a complete method. As a referee, I would send it out, not desk-reject it, because the research direction is timely and the authors are the right people to pursue it. The main revision requests would be: fix the malformed example, and say something substantive about how the two metamodels meet, even if it is only a sketch of possible mapping rules.","headline":"An honest, well-scoped vision paper: the AMDiRE-PerSpecML pairing is a plausible research direction, and the paper's main weakness is that its only illustrative mapping is both hand-crafted and contains a malformed example.","tokens_in":659,"tokens_out":706,"would_cite":false,"duration_ms":22397,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This vision paper claims that integrating AMDiRE's artefact-based structure with PerSpecML's concern-driven guidance gives a concrete pathway from high-level trustworthy AI principles to structured, traceable requirement artefacts.","keywords":["trustworthy AI","requirements engineering","AI-enabled systems","artefact-based requirements engineering","perspective-based requirements elicitation","machine learning systems","EU AI Act","traceability"],"falsifier":"A concrete test would be to take all 60 PerSpecML concerns plus the EU AI Act's seven AI-HLEG requirements and attempt to assign each to an AMDiRE artefact type with a well-formed example requirement. If even one concern admits no artefact type, or if the resulting example entries are malformed or lose traceability to the original stakeholder concern, the central bridging claim would fail.","tokens_in":5730,"feed_emoji":"🤖","tokens_out":6237,"duration_ms":65026,"temperature":0.7,"pith_summary":"The paper is a vision proposal, not a finished method. It claims that two existing requirements engineering approaches can be combined to make trustworthy AI goals operational: AMDiRE supplies structured artefacts with defined content, roles, and milestones, while PerSpecML supplies a catalogue of 60 concerns across 28 ML tasks from five perspectives. The combination would let a team start from a high-level commitment like 'the system must be compliant, ethical, robust' and end with documented requirement artefacts that trace back to stakeholder concerns and regulatory sources. The authors illustrate the idea on an AI hiring platform, mapping concerns such as fairness, explainability, and regulatory compliance to AMDiRE artefacts. If the pathway works, it would give practitioners concrete guidance rather than principles alone, and would make AI trustworthiness auditable during development.","feed_headline":"Two methods, one path to concrete AI trustworthiness requirements","feed_subtitle":"The proposed integration turns high-level ethics and regulation goals into documented, traceable requirement artefacts.","key_machinery":"The load-bearing mechanism is the Perspective-based ML Task and Concern Diagram, PerSpecML's central object that links 60 concerns to 28 ML tasks across five perspectives (system objectives, user experience, infrastructure, model, data). The paper proposes to extend that diagram with regulatory requirements and then attach each concern to one of AMDiRE's artefact content structures — context specification, requirements specification, system specification — plus the role model that assigns validation duties. That attachment is what turns an abstract concern such as explainability into a documented, traceable requirement entry. The illustrative mapping tables in the paper show the intended shape of the mechanism, not its validated behaviour.","core_discovery":"The paper's central claim is that the gap between high-level trustworthiness principles and concrete RE artefacts can be closed by a layered integration. PerSpecML's perspective-based task and concern diagram is the elicitation front end; AMDiRE's artefact, role, and process models are the documentation back end; regulatory frameworks such as the EU AI Act enter as first-class stakeholders whose constraints extend the concern catalogue. The discovery is the bridging move itself: each trustworthiness concern is assigned to an artefact type — context, requirements, or system specification — whose predefined content structure then guides the writing of the requirement. The authors present the mapping as a research direction, scaffolded by an illustrative example rather than a validated method.","pith_inferences":["Modeling regulators as first-class stakeholders could generalize to other normative sources, such as safety standards or data-protection authorities, making the pathway a generic compliance-to-specification channel.","A natural next test is to compare teams using the integrated templates against teams using free-form checklists for the same regulation, measuring completeness, consistency, and traceability of trustworthiness requirements.","If the concern-to-artefact mapping is made explicit and tool-supported, the framework could also support automated consistency checks between regulatory clauses and system specifications.","The authors scope the vision to machine learning; the same structure could extend to generative AI by adding new tasks and concerns for emergent behaviours such as hallucination or misuse."],"forward_implications":["Practitioners would get a defined route from a trustworthiness principle to a documentable requirement artefact, with roles and approval milestones supplied by AMDiRE's process and role models.","Regulatory constraints could be treated as elicited concerns from first-class stakeholders, so compliance evidence would be traceable from EU AI Act clauses to system specifications.","The integration would let teams reason about trade-offs among concerns, such as explainability versus model complexity, during elicitation rather than after implementation.","Tool support could guide concern navigation and artefact generation by extending the existing AMDiRE tool environment.","If validated in practice, the framework would complement existing responsible-AI frameworks by adding a concrete mechanism for specifying trustworthiness requirements."],"supporting_citations":[{"why":"Supplies AMDiRE, the artefact-centric structure with content templates, roles, and milestones that the framework reuses as its documentation backbone.","marker":"[11]"},{"why":"Supplies PerSpecML, the perspective-based concern catalogue (60 concerns, 28 tasks, five perspectives) that drives elicitation.","marker":"[12]"},{"why":"Provides the regulatory requirements of the EU AI Act that the authors propose to fold into the extended concern catalogue.","marker":"[1]"},{"why":"Defines the three-part notion of trustworthy AI (lawful, ethical, robust) and the seven requirements that the paper takes as its normative foundation.","marker":"[13]"},{"why":"Documents practitioners' need for actionable guidelines, motivating the gap the framework addresses.","marker":"[10]"}],"fun_headline_variants":["Closing the gap between AI ethics and requirement specs","A framework to turn AI trust principles into specs","Bridging principles and artefacts for trustworthy AI","Operationalizing AI trust requirements via two-method blend","From trust concerns to concrete requirement artefacts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that PerSpecML's concern catalogue can be extended with regulatory requirements and mapped cleanly onto AMDiRE's artefact content structures at compatible granularity; the paper does not yet show that such a mapping can be systematic rather than hand-crafted.","fun_headline_variants_meta":{"raw":{"variants":["Closing the gap between AI ethics and requirement specs","A framework to turn AI trust principles into specs","Bridging principles and artefacts for trustworthy AI","Operationalizing AI trust requirements via two-method blend","From trust concerns to concrete requirement artefacts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000164,"raw_usage":{"total_tokens":1199,"prompt_tokens":852,"completion_tokens":347,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":468,"completion_tokens_details":{"reasoning_tokens":276}},"tokens_in":468,"tokens_out":347,"duration_ms":3770,"temperature":1.0,"reasoning_tokens":276,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:36:53.321959+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test would be to take all 60 PerSpecML concerns plus the EU AI Act's seven AI-HLEG requirements and attempt to assign each to an AMDiRE artefact type with a well-formed example requirement. If even one concern admits no artefact type, or if the resulting example entries are malformed or lose traceability to the original stakeholder concern, the central bridging claim would fail.","supporting_citations":[{"cited_title":"Artefact-based require- ments engineering: the amdire approach,","cited_arxiv_id":null,"evidence_quote":"Supplies AMDiRE, the artefact-centric structure with content templates, roles, and milestones that the framework reuses as its documentation backbone."},{"cited_title":"Identify- ing concerns when specifying machine learning-enabled systems: A perspective-based approach,","cited_arxiv_id":null,"evidence_quote":"Supplies PerSpecML, the perspective-based concern catalogue (60 concerns, 28 tasks, five perspectives) that drives elicitation."},{"cited_title":"Regulation (eu) 2024/1689 of the european parliament and of the council of 13 june 2024 laying down harmonised rules on artificial intelligence,","cited_arxiv_id":null,"evidence_quote":"Provides the regulatory requirements of the EU AI Act that the authors propose to fold into the extended concern catalogue."},{"cited_title":"Ethics guidelines for trustworthy ai,","cited_arxiv_id":null,"evidence_quote":"Defines the three-part notion of trustworthy AI (lawful, ethical, robust) and the seven requirements that the paper takes as its normative foundation."},{"cited_title":"Trustworthy ai in practice: an analysis of practitioners’ needs and challenges,","cited_arxiv_id":null,"evidence_quote":"Documents practitioners' need for actionable guidelines, motivating the gap the framework addresses."}],"review_version":1}