{"id":"dd45621a-17b8-4790-a3da-bd1c963f53f7","arxiv_id":"2505.07553","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A case study of a maritime company shows that RAG system requirements are elicited iteratively with users, yielding a five-stage process model centered on retrieval requirements.","lead":"This paper reports on how data scientists at a maritime company elicited requirements for a retrieval augmented generation (RAG) system by experimenting with users and building filters. It proposes a five-stage iterative process model for discovering 'retrieval requirements' in expert domains.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim rests on user judgments as valid ground truth, which the paper's own cited literature challenges; no independent validation is provided.","rationale":"The reader's weakest assumption matches my own: the validity and sufficiency of user judgments as ground truth is load-bearing for the central claim. I agree with the reader's conditional verdict; the concern strengthens the need for conditions but does not reject the paper, which is transparent about its limitations and explicitly framed as preliminary. The paper's own citation of Lebovitz et al. makes the tension concrete, and the unmeasured system success means the process model's efficacy is untested. A follow-up validation of user-judged correctness would settle whether the premise holds.","tokens_in":9035,"tokens_out":2828,"duration_ms":30430,"concrete_test":"Re-examine the raw observation/interview data for instances where case handlers' correctness judgments diverged from an independent assessment, such as the actual rule text or a blind expert panel. If divergences appear and are not reconciled, the ground-truth assumption is unsupported. Alternatively, run a quantitative validation at Marcomp: sample filtered RAG outputs, have case handlers rate correctness, and compare with a blind regulatory-expert panel; low agreement would undermine the process model's foundation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The process model's core premise is that case handlers' correctness judgments are valid and sufficient ground truth for eliciting retrieval requirements (Section 4.1: 'allowing them to judge the correctness of the answers'; Section 5: 'because they are the ones who can determine correctness'). The paper cites Lebovitz et al. [13], 'Is AI ground truth really true?', which argues that experts' know-what can be biased, outdated, or inconsistent. The paper never reconciles this tension. In the case, user judgments are the sole evaluation mechanism; there is no independent check that the filters actually improve correctness, and Section 6 admits the system's effects have not been measured. If user judgments are biased—for example, influenced by the LLM hype described in Section 4.4—the elicited 'retrieval requirements' may encode false correctness criteria, and the proposed five-stage model would not reliably produce correct RAG systems. This is an internal tension with the cited literature, not merely an external generalizability concern. The construct-validity limitation (Section 6) acknowledges terminology ambiguity but does not address the accuracy of user judgments as ground truth.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a qualitative case study at a maritime company (Marcomp) that developed a RAG system over a knowledge base of 500,000 historical case-handler answers. Through observation and interviews, the authors describe how data scientists iteratively experimented with case handlers to identify what they call 'retrieval requirements,' eventually building filtering functions that let users constrain retrieval by year and ship characteristics. The paper proposes a five-stage iterative process model: knowledge modeling and experimentation, retrieval strategy, retrievable data management, monitoring and operation, and continuous expectation management. It claims that eliciting retrieval requirements is essential because users are the only ones who can determine correctness.","tokens_in":9198,"tokens_out":3870,"duration_ms":37878,"significance":"If the process model is accepted as a descriptive account, it is a useful early empirical contribution to requirements engineering for RAG systems, a topic with little industry evidence. The study's strengths are its substantial fieldwork (28 hours of observation plus multiple interviews), its transparent reporting of limitations, and its grounding in practitioner quotes and observed events. The proposed distinction between retrieval strategy and retrievable data management is a plausible and potentially reusable conceptual contribution. However, the normative force of the claims currently exceeds what the single-case evidence can support, because outcome data are absent and the reliability of user correctness judgments is not reconciled with the paper's own cited literature.","major_comments":[{"comment":"The central claim that data scientists must rely on case handlers because 'they are the ones who can determine correctness' is in tension with the paper's own citation of Lebovitz et al. [13], which documents that experts' 'know-what' can be biased, outdated, or inconsistent. The manuscript provides no independent validation of the case handlers' judgments, and Section 6 concedes that the system's effects have not been measured. As written, the evidence supports only that user judgments were used as a pragmatic proxy in this case, not that they are valid ground truth. Please either soften the normative claim or show how the proposed process would detect or correct biased user judgments.","section":"§4.1, §5, Abstract"},{"comment":"The conclusion that eliciting 'retrieval requirements' is essential to ensure output correctness is not supported without outcome measurement. The authors state in Section 6 that they 'have not yet measured the effects of the RAG system' and 'cannot determine whether its development has been a success.' Without any outcome measure, the paper cannot establish that the five-stage process actually produces correct RAG outputs; it can only describe observed practice. The prescriptive conclusions should be reframed as a hypothesis or as a descriptive account of current practice until outcome data are collected.","section":"§5, §6"},{"comment":"The internal structure of the proposed process model is presented inconsistently. The text in Section 5 says the authors 'identify four sequential steps' with expectation management as an 'underlying continuous activity,' while Figure 2 and the conclusion describe a 'five-stage iterative process.' Since the process model is the paper's main contribution, the relationship between expectation management and the other four stages must be clarified: is expectation management a stage, a cross-cutting activity, or both? This distinction affects how the model should be applied and evaluated.","section":"§5, Figure 2"}],"minor_comments":[{"comment":"Table 1 does not specify the number of interviews conducted with the Data Architect versus the AI Solution Engineer; reporting the exact counts and timing of each interview round would aid reproducibility and transparency.","section":"Table 1"},{"comment":"The description of the analysis would benefit from reporting whether the temporal-bracketing phase boundaries were validated with participants (e.g., member checking), since the phase structure is central to the proposed model.","section":"§3"},{"comment":"There is a typographical error in the quote attribution: 'Architecht' should be 'Architect.' The same misspelling appears in the later quote from Magnus.","section":"§4.2"},{"comment":"The heading 'TOW ARDS REQUIREMENTS ENGINEERING FOR RAG' appears to contain a line-break artifact; it should read 'TOWARDS REQUIREMENTS ENGINEERING FOR RAG.'","section":"§5 heading"},{"comment":"Reference [17] (Runeson and Höst) gives only the starting page '131'; it should include the full page range and, ideally, a DOI.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a legitimate short-paper candidate: the descriptive field data are valuable, and the authors are unusually candid about limitations. The main concern is that the abstract and conclusions overstate what a single case with no outcome measurement can support, especially on the validity of user correctness judgments. The revision route should focus on recalibrating the claims rather than adding new data."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick take: this is a well-written, transparent case study from a maritime firm, and it makes a modest but real contribution: it gives us an empirically grounded five-stage process for eliciting 'retrieval requirements' in RAG systems. If you work on RE4AI, it's worth reading.\n\nWhat's new: the paper extends Amershi and Vogelsang/Borg to RAG, and the two middle stages—retrieval strategy and retrievable data management—are the most distinct. The filtering-function story is concrete and shows the shift from adding data to constraining retrieval. That's a nice observation.\n\nWhat's good: the authors are unusually honest. They say generalization is not established, construct validity is a problem, and they haven't measured whether the system works. They share quotes and observations so you can see the chain. No sign of overclaiming.\n\nSoft spots: the load-bearing premise is that case handlers can determine correctness. The paper explicitly states that, and it cites Lebovitz et al. on expert know-what being biased or outdated. It never reconciles those. If user judgments are skewed—say, by the LLM hype the paper describes—then the derived requirements might be wrong, and the process might not produce correct systems. That's an internal tension, not just an external generalizability concern. The remedy is to acknowledge it more directly and, ideally, add some independent check on output quality.\n\nSecond, the model is inferred from one case with no outcome data. That's a limitation the authors admit, but it means the five stages are a hypothesis, not a validated framework. The 'four sequential steps' vs 'five-stage' wording is harmless.\n\nThe construct validity issue is acknowledged. I'd have liked a bit more detail on how they validated terminological interpretations, but it's acceptable for a short paper.\n\nBottom line: this is a legitimate starting point for RE of RAG systems. It deserves a serious referee and publication as a preliminary study; it should not be treated as a validated process. I'd read it and cite it.","headline":"A clear, honest single-case study that proposes a five-stage RE process for RAG systems; the model is plausible but rests entirely on user-provided correctness judgments, and the paper never squares that with the literature it cites.","tokens_in":9736,"tokens_out":1396,"would_cite":true,"duration_ms":13150,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that retrieval requirements for RAG systems are elicited through iterative experimentation with expert users, and proposes a five-stage process model for doing so.","keywords":["Requirements Engineering","Retrieval Augmented Generation","RAG","GenAI","RE4AI","case study","maritime industry","retrieval requirements"],"falsifier":"If a comparable RAG deployment in another expert organization achieved equally correct outputs (as judged by independent experts) using pre-specified requirements and automated retrieval evaluation without iterative user experimentation or user-controlled filters, the claimed necessity of the five-stage user-experimentation process would be refuted.","tokens_in":8819,"feed_emoji":"🤖","tokens_out":7846,"duration_ms":67345,"temperature":0.7,"pith_summary":"This short paper argues that requirements for retrieval-augmented generation (RAG) systems—where an LLM pulls relevant stored documents or past answers into its prompt before answering—cannot be specified up front. Through a case study of a maritime service provider building a RAG system over 500,000 previously written expert answers, it shows that data scientists had to discover 'retrieval requirements' through iterative experimentation with the expert users who judge whether generated answers are correct. The paper presents an empirically grounded, five-stage iterative process model: knowledge modeling and experimentation, retrieval strategy, retrievable data management, monitoring and operation, and continuous expectation management. The contribution matters because RAG is a pragmatic way to integrate LLMs into organizations, yet software and requirements engineering research has lacked concrete industry guidance on how such systems' requirements are actually elicited.","feed_headline":"RAG requirements emerge from user experiments, case study shows","feed_subtitle":"Expert users judge correctness, so retrieval requirements are discovered iteratively, not specified upfront.","key_machinery":"The central object is the 'retrieval requirement' (RR), defined in the model as a requirement about what should be retrievable from the knowledge base as input to the LLM. The mechanism carrying the argument is the iterative experimentation loop: data scientists release an imperfect RAG system, expert users judge output correctness, those judgments reveal gaps in retrievable data, and each gap is addressed by a compensation such as a filtering function or mandatory human review. That loop is organized by the paper's five-stage process model, which gives requirements engineering a concrete handle on a system class where the traditional 'specify first' approach breaks down.","core_discovery":"On the paper's own terms, the central discovery is that RAG system correctness is not a property the data science team can define alone; it resides with experienced case handlers, who know current rules, strategic interpretation norms, and the infinite combinations of vessel characteristics that make context unique. Because the knowledge base of past answers cannot cover every context and becomes outdated as rules change, the data scientists repeatedly hit retrieval failures and compensated by giving users filtering controls (by year, vessel type, nationality, age, and similar factors) and by requiring a human check of every generated answer. Each such compensation is a 'retrieval requirement'—a decision about which parts of the knowledge base may legitimately be retrieved as input to the LLM. The paper generalizes these observations into a five-stage iterative process model in which requirements emerge from use, are refined through monitoring, and unfold under continuous—and in this case partly abandoned—expectation management.","pith_inferences":["Beyond the paper: the same pattern—legacy corpus, dated rules, and context combinations too numerous to encode—is likely to appear in legal, medical, and financial RAG deployments, so the five stages may transfer across expert domains.","Beyond the paper: because correctness is delegated to expert users, the process inherits their blind spots; a systematic error shared by experts would be encoded into the retrieval requirements, suggesting organizations should add independent outcome checks.","Beyond the paper: the filter choices case handlers make could be logged and used to learn which context features predict retrieval relevance, potentially automating parts of the retrieval strategy stage in later deployments.","Beyond the paper: the observation that expectation management was abandoned implies the real constraint on RAG adoption may be organizational—managing hype—rather than technical retrieval quality."],"forward_implications":["RAG requirements engineering should treat user experimentation, not upfront specification, as the primary discovery mechanism.","Retrieval strategy and retrievable data management become first-class RE stages, because runtime control of what the LLM can access substitutes for retraining in traditional ML.","Monitoring live system use is itself a requirements elicitation activity: new retrieval requirements surface only after real users interact with early versions.","Expectation management must be planned as an ongoing activity, but the case shows it can fail under LLM hype, leaving users with unmet expectations even as the system improves.","The human check on every generated answer is a deliberate requirements decision, not just a safety feature, because interpretation norms and strategic motives cannot be captured in the corpus."],"supporting_citations":[{"why":"Defines retrieval-augmented generation and supplies the four-step pipeline (indexing, retrieval, augmentation, generation) that the case system instantiates.","marker":"[15]"},{"why":"Catalogues common RAG failure points such as missing content and extraction failures, which the paper's retrieval strategy and retrievable data management stages address.","marker":"[4]"},{"why":"Provides the software-engineering-for-ML case-study baseline whose workflow stages, including monitoring, the proposed RAG process extends.","marker":"[2]"},{"why":"Documents how data scientists approach requirements for ML systems and argues that knowledge-base diversity matters, a premise behind the knowledge modeling stage.","marker":"[26]"},{"why":"Supplies evidence that defining correctness criteria is a core difficulty in ML systems engineering, motivating the shift to user-judged correctness.","marker":"[9]"},{"why":"Provides the case-study research guidelines that structure the empirical method.","marker":"[17]"},{"why":"Supplies temporal bracketing as the analysis strategy used to identify the five process stages.","marker":"[12]"},{"why":"Supplies the grounded theory approach used for initial coding and category emergence.","marker":"[20]"}],"fun_headline_variants":["RAG correctness is user-defined, not predetermined","Retrieval requirements emerge from user feedback, not specs","For RAG, requirements are discovered, not specified","Expert users define RAG retrieval correctness","Case study: RAG requirements need iterative user testing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole process model depends on the assumption that the experienced users' judgments of whether an answer is correct are a valid and sufficient ground truth for deciding what the system should retrieve.","fun_headline_variants_meta":{"raw":{"variants":["RAG correctness is user-defined, not predetermined","Retrieval requirements emerge from user feedback, not specs","For RAG, requirements are discovered, not specified","Expert users define RAG retrieval correctness","Case study: RAG requirements need iterative user testing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000305,"raw_usage":{"total_tokens":1699,"prompt_tokens":842,"completion_tokens":857,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":458,"completion_tokens_details":{"reasoning_tokens":784}},"tokens_in":458,"tokens_out":857,"duration_ms":8074,"temperature":1.0,"reasoning_tokens":784,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:13:18.896569+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If a comparable RAG deployment in another expert organization achieved equally correct outputs (as judged by independent experts) using pre-specified requirements and automated retrieval evaluation without iterative user experimentation or user-controlled filters, the claimed necessity of the five-stage user-experimentation process would be refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines retrieval-augmented generation and supplies the four-step pipeline (indexing, retrieval, augmentation, generation) that the case system instantiates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Catalogues common RAG failure points such as missing content and extraction failures, which the paper's retrieval strategy and retrievable data management stages address."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the software-engineering-for-ML case-study baseline whose workflow stages, including monitoring, the proposed RAG process extends."},{"cited_title":"and Borg, M","cited_arxiv_id":null,"evidence_quote":"Documents how data scientists approach requirements for ML systems and argues that knowledge-base diversity matters, a premise behind the knowledge modeling stage."},{"cited_title":"and Yoshioka, N","cited_arxiv_id":null,"evidence_quote":"Supplies evidence that defining correctness criteria is a core difficulty in ML systems engineering, motivating the shift to user-judged correctness."},{"cited_title":"and Höst, M","cited_arxiv_id":null,"evidence_quote":"Provides the case-study research guidelines that structure the empirical method."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies temporal bracketing as the analysis strategy used to identify the five process stages."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the grounded theory approach used for initial coding and category emergence."}],"review_version":1}