{"id":"099f3310-3460-40cf-9c98-e94fbe6df656","arxiv_id":"2411.19554","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A six-user qualitative pilot shows a custom GPT RAG chatbot for university students is liked for tone and structure but suffers from inaccuracy, omitted information, and broken links.","lead":"This pilot paper describes a ChatGPT-based RAG chatbot built for University of Milano-Bicocca students, and reports qualitative feedback from six testers. It finds the chatbot was seen as friendly and fast but unreliable due to hallucinations and broken links.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Accuracy and 'often neglects relevant information' findings rest on unquantified monitoring of six free-form sessions; no coding scheme, query log, or inter-rater check supports the frequency claims.","rationale":"The reader's weakest-assumption analysis focused on generalizability from six mostly Psychology students to the whole Unimib population. That is a legitimate concern, and the authors themselves flag it in Section 4. My stress-test identifies a second, more internal threat: the paper's frequency claims ('often,' 'sometimes,' 'not always') are supported only by free-form interview impressions and unspecified monitoring, without a reproducible evaluation protocol. This is not a reason to reject the paper—it is an honest pilot study whose qualitative descriptions are plausible—but it does mean the contribution list in Section 5, especially items 7 and 8, overstates what the current data can establish. The reader's conditional verdict already captures the need for stronger evidence; my concern reinforces that condition by adding a concrete methodological requirement: a fixed benchmark with gold answers and independent annotation. I therefore recommend keeping the verdict unchanged while making the condition more specific. I chose 'partial' agreement because I agree with the reader that the sample limits generalizability, but my primary concern is the internal validity of the error-frequency claims rather than the sampling frame alone. No ad hominem is intended; the issue is with the evidence base, not the authors' conduct.","tokens_in":9000,"tokens_out":3593,"duration_ms":36060,"concrete_test":"Reconstruct the exact prompt, uploaded documents, and link list described in Section 3.3, then build a fixed benchmark of 20 questions with gold answers extracted from those same documents (e.g., thesis deadlines, scholarship ISEE procedure, canteen prices). Pose each question to the same Custom GPT configuration, record the raw responses, and have two annotators who were not involved in the design independently code each response for factual correctness, completeness (omission of gold facts relevant to the question), and URL validity (does the link resolve to a real Unimib page?). Compute per-item accuracy, omission rate, link validity rate, and Cohen's kappa between annotators.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central contribution is the claim that the chatbot's limitations—occasional inaccuracy, frequent omission of relevant information from uploaded documents, and broken links—'severely impacted the overall experience' (Abstract, Sections 3.5 and 3.6). For this claim to hold, the evaluation must reliably measure those failures. Section 3.4 says the authors 'closely monitored each query to ensure that GPT's responses were accurate and aligned with the information contained in the uploaded documents,' but no monitoring criteria, error taxonomy, query log, or inter-rater reliability check is reported. The qualitative usability test used free interaction (Section 3.4), so different participants likely asked different questions; this makes it impossible to know whether 'often' reflects systematic RAG behavior or a few idiosyncratic queries. Section 3.5 reports 'not always 100% accurate,' 'sometimes provided nonexistent or broken links,' and 'often neglect to report relevant information' without counts, examples, or a denominator. Section 4 acknowledges the small sample size as a limitation, but it does not acknowledge that the measurement of the headline errors is unstandardized. Section 5 lists 'Evaluation of user interaction and LLM + RAG reliability impact' as a contribution, and Section 3.6 motivates the redesign by calling hallucination and broken links 'glaring issues'—but that redesign is justified entirely by the unquantified observations. The core correctness-risk is internal validity: the evidence supports 'some users encountered some errors,' not 'often neglected relevant information' or 'severely impacted experience.' This does not contradict the authors' honesty or the plausibility of the findings, but it makes the generalizable claims fragile.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Unimib Assistant, a RAG-based chatbot built with OpenAI's Custom GPTs for University of Milano-Bicocca students. A needfinding phase with six semi-structured interviews informed the first prototype; a qualitative usability test with six other students evaluated it. The authors report that the chatbot was perceived as fast, ease-to-use, and friendly, but that it also produced inaccurate information, neglected relevant information present in uploaded documents, and sometimes generated broken links. A redesign phase addressed these issues by adding links, merging documents, and changing the chatbot's name and logo. The paper closes with a list of limitations and future plans, including a larger usability test and API integration.","tokens_in":9213,"tokens_out":3777,"duration_ms":33726,"significance":"The paper makes a practical contribution by showing how a low-cost, non-programmer-friendly RAG chatbot can be built and iteratively refined for a university context. Its strengths include the user-centered design process, the separation of needfinding and evaluation participant groups, and an honest, explicit list of system limitations. The qualitative findings align with known LLM reliability issues and could be useful guidance for similar pilots. However, the significance is limited by the lack of systematic analysis of the evaluation data: the central claims about the frequency and nature of the chatbot's errors are not backed by quantified evidence or replicable coding procedures. The paper would be more valuable if it either provided such evidence or carefully restricted its claims to anecdotal observations.","major_comments":[{"comment":"The central negative findings—'not always 100% accurate,' 'often neglect to report relevant information,' and 'sometimes provided nonexistent or broken links'—are based on unstandardized observation of six free-form testing sessions. No query log, error taxonomy, counts, or inter-rater reliability check is reported. Because participants interacted freely with the system, they likely asked different questions, so the frequency words 'often' and 'sometimes' cannot be interpreted as properties of the system rather than of a few idiosyncratic queries. This measurement gap underlies the abstract's claims, the redesign rationale in §3.6, and contribution item 8 in §5.","section":"Abstract, §3.4–§3.5, §3.6"},{"comment":"The statement that the authors 'closely monitored each query to ensure that GPT's responses were accurate and aligned with the information contained in the uploaded documents' presupposes a verification procedure that is not described. There is no operational definition of 'accurate' or 'aligned,' no indication of whether all queries or only a subset were monitored, and no report of how many responses were judged to contain errors. This is load-bearing because the paper's main negative conclusions stem from this monitoring.","section":"§3.4"},{"comment":"The results are presented as a bullet list without any participant-level detail, direct quotes, or counts. It is thus impossible to know whether each finding reflects one participant, a majority, or all six. For example, 'The chatbot was considered fast, easy to use and understand, and reliable when it provided adequate sources' is reported as a single result, but the reader cannot tell how many testers expressed this view. This undercuts the strength of both the positive and negative conclusions.","section":"§3.5"},{"comment":"The limitations section correctly notes the small sample size but does not acknowledge that the measurement of the headline errors is also unstandardized. Table 1 lists 'Usability Testing / Small sample size affecting generalizability,' but omits the more fundamental issue that the frequency claims were not derived from a systematic analysis of recorded sessions or logs. The planned larger study should include a defined error taxonomy and inter-rater validation for the accuracy and source-link checks.","section":"§4, Table 1"}],"minor_comments":[{"comment":"The demographic information is reported for each participant group, but the recruitment is described only as 'university colleagues' invited through personal contacts; the paper does not discuss how this may have influenced participants' familiarity with the research team or their motivation.","section":"§3.1 and §3.4"},{"comment":"The prompt and links are shown as screenshots (Figures 1–2), but the text does not provide a verbatim transcript or structured summary. This makes it hard for other researchers to replicate the exact configuration or assess how the prompt may have influenced the observed behavior.","section":"§3.3"},{"comment":"The bullet list mixes researcher observations with participant reports without distinction. For example, 'The information provided by the chatbot was not always 100% accurate' appears to be a researcher judgment, while 'The friendly tone... was considered good' is a participant report; distinguishing these would improve clarity.","section":"§3.5"},{"comment":"There are several formatting and minor textual issues: reference [4] contains 'Learning edidattica online' in the URL, reference [22] has an extra space in the URL, and the citation numbering in §1.1 is not sequential ([12], then [16], then [13], [14], [15], [5], [21], [22]).","section":"References"},{"comment":"The term 'needfinding' is written inconsistently as 'need-finding,' 'need finding,' and 'needfinding'; please standardize.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper describes a useful pilot study and the authors are transparent about many limitations. My main concern is the gap between the strength of the claims about the chatbot's reliability problems and the evidence: six free-form sessions with no systematic coding or error measurement. This can be addressed by reanalyzing existing data if session logs or recordings exist, or by explicitly reframing the findings as anecdotal and hypotheses-generating. The topic fits a HCI/educational technology venue, but the empirical backbone needs to be strengthened before I could recommend acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a straightforward, honest pilot study: a custom GPT with RAG for University of Milano-Bicocca students, built from a six-person needfinding phase and evaluated with six other students in a qualitative usability test. The main contribution is descriptive: a concrete account of what happens when a non-expert team builds a RAG chatbot on the OpenAI custom GPTs platform. The practical constraints they document—20-document upload limit, unclickable links, premium paywall, inability to inspect referenced documents—are genuinely useful for other universities considering the same route. The redesign phase, triggered by user feedback on links and branding, is a sensible user-centered loop. The writing is clear, the limitations section is candid, and the authors do not oversell the system's maturity.\n\nThe soft spots are the ones the stress-test note flags. The claim that the chatbot 'often neglected to report relevant information' and that these failures 'severely impacted the overall experience' is not backed by any measurement. The usability test used free interaction with no query log, no error taxonomy, no inter-rater check, and no counts or examples. Different users asked different questions, so 'often' has no clear denominator. A reader can only conclude that some users encountered some errors. That is a real internal-validity problem, and it matters because the redesign is justified by those unquantified observations. The small, homogeneous sample (five of six from Psychology) is acknowledged, and the lack of a systematic qualitative analysis (no coding scheme, no quotes) also limits the strength of the claims.\n\nStill, the paper does not pretend to be a controlled experiment. It is an honest pilot, and the central message—students find such a chatbot friendly and useful, but its factual reliability is questionable—is credible and consistent with prior RAG/LLM literature the authors cite. The contribution list in Section 5 is overreaching for a 6-user pilot, but that is a minor framing issue, not a fatal flaw.\n\nThis is a workshop-level paper with a clear, reproducible design process and a useful limitations table. It deserves a serious referee, especially one who can push the authors to either quantify the error claims or soften them to 'some participants reported occasional inaccuracies.' I would accept it conditionally, with the main revision being to align the claims with the evidence. I would not cite it in my own work beyond a footnote example of user-reported RAG limitations.","headline":"Honest, small-scale pilot of a RAG chatbot for university students; the qualitative findings are plausible but the headline frequency claims outrun the evidence.","tokens_in":9835,"tokens_out":603,"would_cite":false,"duration_ms":7542,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This pilot study reports that a RAG-based chatbot built with OpenAI's Custom GPTs was perceived by its six student testers as fast, friendly, and reliable when sources worked, but that inaccuracies, omitted document content, and broken…","keywords":["ChatGPT","Retrieval-Augmented Generation","student-friendly chatbot","user experience","human-centered computing","usability testing","question answering","hallucination"],"falsifier":"A concrete test is to run the redesigned assistant with a larger, department-balanced and international sample on fixed guided tasks, then independently check every answer against the uploaded documents for factual accuracy and completeness and check every returned link's HTTP status; if error rates remain high, the pilot's accuracy and trust problems are systematic, and if they approach zero, those problems were artifacts of the small sample and early prototype.","tokens_in":8782,"feed_emoji":"🎓","tokens_out":8703,"duration_ms":74882,"temperature":0.7,"pith_summary":"This pilot study reports the design and first usability test of Unimib Assistant, a question-answering chatbot for students of the University of Milano-Bicocca that combines GPT-4 with retrieval-augmented generation (RAG) through OpenAI's Custom GPTs feature. The authors claim that the six students who evaluated it experienced the assistant as fast, easy to use, well structured, and trustworthy when it supplied working sources, but that occasional inaccurate answers, omissions of information present in the uploaded documents, and unclickable links significantly damaged satisfaction and trust. The work's purpose is to show that such an assistant can be configured without programming, grounded in a user-centered cycle of needfinding interviews, prototype refinement, and qualitative usability testing. A sympathetic reading takes the contribution to be a replicable design pattern plus a candid catalogue of where a RAG chatbot on this platform still fails, rather than a statistical demonstration of accuracy.","feed_headline":"University chatbot pilot: fast, friendly, sometimes wrong","feed_subtitle":"Six testers liked its tone and speed; inaccuracies and broken links hurt trust.","key_machinery":"The central object is Unimib Assistant itself, a custom GPT built on the OpenAI Custom GPTs feature, configured manually with a purpose-and-behavior prompt, a set of university-related links, and uploaded PDF and PowerPoint documents. The mechanism that carries the argument is retrieval-augmented generation (RAG): when a student asks a question, GPT-4 generates an answer using both the user's query and the information retrieved from those uploaded sources, which is what lets a generic LLM speak about specific university procedures such as thesis requirements, scholarships, and Erasmus calls. The evaluation mechanism is the user-centered design loop, with needfinding interviews to select content, a qualitative usability test using think-aloud interaction and a semi-structured interview, and a redesign phase, because it is this loop that produces the reported strengths, weaknesses, and fixes. The paper treats this prompt-plus-RAG-plus-feedback configuration as the reusable core that other universities could copy.","core_discovery":"On the paper's own terms, the central discovery is that a GPT-based RAG chatbot can meet students' desire for quick, direct, conversational answers, and that its perceived reliability hinges on the availability of correct, clickable sources. In usability tests with six Unimib students, the chatbot's simple ChatGPT interface, friendly tone, and structured answers were judged positively, and users called it fast, easy to use, and reliable when it provided adequate sources. At the same time, the system did not always return fully accurate information, sometimes ignored relevant content that was present in the uploaded prompt and documents, and occasionally fabricated broken links; the authors state that these issues severely impacted the overall experience and undermined trustworthiness. The authors respond with a redesign that adds more links, merges uploaded documents, and changes the name and logo, and they argue the whole recipe is easy for other institutions to replicate.","pith_inferences":["I would infer that the decisive design variable is not retrieval quality alone but what the model does when retrieval fails: the fabricated links suggest the assistant prefers to answer over admitting it lacks a source, so a retrieval-confidence threshold or an explicit 'I don't know' fallback would likely be the next effective intervention.","I would infer that the 'neglected relevant information' failures could be measured objectively by prompting the same questions repeatedly with and without specific source documents and comparing how often the model cites each uploaded fact; random variations would indicate retrieval instability rather than systematic misunderstanding.","A testable extension is to monitor live usage with logging of queries, returned links, and HTTP status codes, so broken-link incidence and omission rates can be tracked over redesigns instead of relying on recollection from six interviews."],"forward_implications":["Other universities can build a similar assistant without custom software by keeping the base prompt structure and swapping in their own links and documents.","Users' trust in such a chatbot depends less on interface aesthetics than on the model's ability to give accurate answers and working source links; a broken link is treated by testers as a reliability failure.","Adding more documents and links and merging files can reduce, but evidently does not eliminate, hallucinations and fabricated links, so institutions should still plan for disclaimers, human support, or both.","The qualitative results are design guidance rather than generalizable evidence; the paper's own planned next step is a larger, multi-department usability test with guided tasks."],"supporting_citations":[{"why":"Defines retrieval-augmented generation, the core technique the chatbot uses to combine retrieval with generation.","marker":"[6]"},{"why":"Documents the RAG and semantic search feature for OpenAI's custom GPTs that implements the assistant's retrieval.","marker":"[17]"},{"why":"Introduces the Custom GPTs feature that is the platform on which Unimib Assistant is built.","marker":"[18]"},{"why":"Surveys hallucination in natural language generation, the failure mode the paper's evaluation centers on.","marker":"[4]"},{"why":"Reports that ChatGPT fabricates a large share of references for medical questions, motivating the paper's focus on link accuracy and source reliability.","marker":"[12]"},{"why":"Shows a QA-RAG chatbot still struggles with external data that conflicts with the model's prior knowledge, a limitation echoed in this study's accuracy problems.","marker":"[5]"},{"why":"Provides the nearest prior case of a RAG chatbot in higher education increasing engagement, which this study extends to a university-wide assistant.","marker":"[14]"}],"fun_headline_variants":["Student chatbot wins on tone, loses on accuracy","Friendly campus chatbot can't always source its answers","RAG chatbot for students: liked, but not always right","Unimib chatbot pilot: fast answers, spotty links","Chatbot wins users over, then fails on facts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that six needfinding interviews, five of them with Psychology master's students, plus six evaluation interviews, can represent the information needs and user experience of the whole Unimib student body, including international students.","fun_headline_variants_meta":{"raw":{"variants":["Student chatbot wins on tone, loses on accuracy","Friendly campus chatbot can't always source its answers","RAG chatbot for students: liked, but not always right","Unimib chatbot pilot: fast answers, spotty links","Chatbot wins users over, then fails on facts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000183,"raw_usage":{"total_tokens":1337,"prompt_tokens":988,"completion_tokens":349,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":604,"completion_tokens_details":{"reasoning_tokens":270}},"tokens_in":604,"tokens_out":349,"duration_ms":3296,"temperature":1.0,"reasoning_tokens":270,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:03:50.191446+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test is to run the redesigned assistant with a larger, department-balanced and international sample on fixed guided tasks, then independently check every answer against the uploaded documents for factual accuracy and completeness and check every returned link's HTTP status; if error rates remain high, the pilot's accuracy and trust problems are systematic, and if they approach zero, those problems were artifacts of the small sample and early prototype.","supporting_citations":[{"cited_title":"(September 2024)","cited_arxiv_id":null,"evidence_quote":"Documents the RAG and semantic search feature for OpenAI's custom GPTs that implements the assistant's retrieval."},{"cited_title":"(November 2023)","cited_arxiv_id":null,"evidence_quote":"Introduces the Custom GPTs feature that is the platform on which Unimib Assistant is built."},{"cited_title":"& Fung, P","cited_arxiv_id":null,"evidence_quote":"Surveys hallucination in natural language generation, the failure mode the paper's evaluation centers on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports that ChatGPT fabricates a large share of references for medical questions, motivating the paper's focus on link accuracy and source reliability."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows a QA-RAG chatbot still struggles with external data that conflicts with the model's prior knowledge, a limitation echoed in this study's accuracy problems."}],"review_version":1}