{"id":"09be1118-7d3f-481c-a377-018b462fa9e3","arxiv_id":"2505.10490","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A planned German university survey will compare trust and usage of a customized university LLM chatbot against ChatGPT; no empirical results are reported yet.","lead":"This paper proposes a field study to test whether university-branded LLM chatbots are trusted and used differently than commercial ChatGPT. It presents hypotheses about trust, hallucinations, privacy, and sustainability, but reports no data yet.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Planned comparison cannot isolate user-salient customizations: Table 2 shows the university LLMaaS differs from ChatGPT in temperature, model access, image generation, data governance, VPN access, token display, and training materials, so observed trust/usage differences would not be attributable…","rationale":"The paper is a prequel with no empirical results, so the reader's UNVERDICTED verdict is appropriate; nothing in this stress-test changes that. The main concern is that even the planned study, as specified in Section 3.2/Table 2, cannot cleanly test the central claim because the university system and ChatGPT differ in multiple substantive, non-cosmetic features. This makes the attribution of any observed effect to user-salient customizations (branding, interface) unsupported by the current design. The reader's weakest assumption correctly identified the Table 2 confounds; I agree fully. The concern is load-bearing because the paper's stated contribution depends on separating user-salient from functional customizations. The paper is otherwise clearly written, hypotheses are grounded in prior work, and the planned multi-channel recruitment and within-user repeated measures have merit. One additional side issue is that reference [46] contains a placeholder URL (example.com/your-thesis-url), so the WärtsiläGPT example lacks evidentiary support, but this is not central. The suggested concrete test would give the authors a feasible way to isolate the branding effect within the same deployment context and would convert the planned field study from a confounded comparison into a test that can actually support or refute H1-H6.","tokens_in":12380,"tokens_out":5815,"duration_ms":60279,"concrete_test":"Add a randomized experimental control condition to the field study: present a subset of participants with an otherwise identical Azure-hosted university LLM in a neutral, unbranded interface but with the same temperature, model set, image generation, data governance, VPN restriction, and token display as the university version, and compare trust/usage ratings against both the branded university version and ChatGPT. If the neutral version yields significantly different trust/usage scores from the branded version and matches ChatGPT, branding is the active customization; if the neutral version tracks the branded version, the Table 2 feature bundle rather than user-salient customization explains the effect.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 1) is that customizations users can see, including purely cosmetic choices such as corporate branding, change trust and usage because users read them as cues about capability, purpose, and process. The planned field study compares the university's LLMaaS to ChatGPT, but Section 3.2/Table 2 lists multiple substantive differences between the two systems: temperature is fixed to 0; multiple models and DALL-E image generation are available; chat data are excluded from training and processed under EU GDPR; access is restricted to internal networks/VPN; token usage is displayed as a percentage; and formal training materials are provided. This contradicts the paper's statement that 'customizations regarding the data and model were kept to a minimum.' If users' trust, hallucination caution, privacy perceptions, or resource-conscious prompting differ, the effect could be driven by any of these features rather than by the user-salient branding/interface customization the paper emphasizes. Because the comparison is observational and participants know which system they are rating, prior experience, organizational trust, and self-selection are additional confounds. The survey instrument does not specify statistical or design controls for these dimensions. Therefore, as designed, the study would not provide a clean test of the central claim even after the planned data collection.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a 'late-breaking' study-design paper. It argues that user-salient customizations of LLM-as-a-Service systems—such as corporate branding and interface changes—can influence end users' trust and usage, even when the customizations have no functional impact. The authors derive six hypotheses (H1–H6) from literature on trust, hallucinations, privacy, and sustainable AI use, and describe a planned cross-sectional field study at a German university that compares the institution's customized ChatGPT-based LLMaaS with commercial ChatGPT. The paper presents no empirical data; its stated purpose is to stimulate discussion and obtain peer feedback on the design.","tokens_in":12744,"tokens_out":4568,"duration_ms":42179,"significance":"The paper's contribution is a structured taxonomy of LLMaaS customizations (Table 1) and a set of theory-derived hypotheses linking customized identity to trust and behavior. If the planned study were cleanly able to isolate branding/interface effects, the results could inform organizational AI adoption. The authors are also transparent that this is a prequel. However, the test they propose is confounded by substantive system differences, so the central claim is not directly estimable from the planned comparison. The intended contribution therefore needs substantial design work rather than minor polishing.","major_comments":[{"comment":"The statement that 'customizations regarding the data and model were kept to a minimum' is contradicted by the table, which lists temperature fixed to 0, access to multiple models, DALL-E image generation, GDPR-compliant processing within the EU, VPN-only access, token usage displayed as a percentage, and formal training materials. These are substantive differences in model behavior, privacy guarantees, and functionality. Any observed differences in trust, caution, privacy perception, or resource-conscious prompting between the campus system and ChatGPT could be driven by these factors rather than by corporate branding or interface design. The planned observational comparison cannot isolate the user-salient customizations that are the paper's central independent variable. Please add explicit measurement or statistical control of these confounds (e.g., perceived data governance, model awareness, feature usage) and discuss the residual attribution threat in the design; alternatively, restructure the study as an experimental manipulation of branding/UI with backend features held constant.","section":"3.2, Table 2"},{"comment":"The newly introduced construct 'AI resource consciousness' (Section 2.5) has no validated measure; the survey uses self-developed items without reported pilot testing, reliability, or validity evidence. The same applies to the self-developed hallucination caution and experienced-hallucination items (Section 3.3). Without at least a pilot psychometric assessment, the planned hypothesis tests (H3, H4, H6) cannot distinguish true effects from measurement artifact. Please include item texts in full and report a validation plan (e.g., factor analysis, internal consistency, test-retest) for the self-developed scales.","section":"3.3"},{"comment":"The sampling plan is underspecified in ways that affect the central comparison. The paper targets N=250 with equal distribution between users of the campus system and ChatGPT, but it does not state how non-users of the campus system are recruited or how self-selection into either group is handled. For participants who use both systems, repeated ratings of both systems create order and carryover effects; no counterbalancing or mixed-model analysis is described. Please provide a concrete recruitment and analysis plan that addresses selection bias and within-subject dependencies.","section":"3.1"}],"minor_comments":[{"comment":"In the paragraph beginning 'Having explored potential impacts', 'LMMaaS' should be 'LLMaaS'.","section":"Section 2.5"},{"comment":"In the first row, 'Temperature1set' is missing a space and should read 'Temperature set'.","section":"Table 2"},{"comment":"Reference [46] contains the placeholder URL 'https://example.com/your-thesis-url' and appears to be an incomplete master's thesis citation; it should be completed.","section":"References"},{"comment":"Figure 1 is referenced but the text does not describe its content sufficiently; ensure the figure is included and legible in the submission.","section":"Section 3.3"},{"comment":"The text states that survey questions are available in the supplementary material; if the supplementary material is not part of the submission, it must be provided for review.","section":"Section 3.3"},{"comment":"H3 predicts less cautious behavior and H4 predicts fewer experienced hallucinations; because less cautious users may also detect fewer hallucinations, the relationship between the two hypotheses needs clarification, distinguishing non-occurrence from non-detection.","section":"Section 2.3"}],"recommendation":"major_revision","confidential_remarks":"This is essentially a study proposal, and its fit with a journal that primarily publishes empirical results is questionable. The confound identified in Table 2 is severe: even with revision, the planned field study may not provide a clean test of the central claim. I would not reject solely because no data are present, given the 'late-breaking' framing, but the design flaw must be addressed, ideally by adding a pilot validation and explicit control/analysis strategies, or by reframing the paper as a design/position paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: this is a late-breaking design paper, not an empirical one. It presents six hypotheses and a planned survey comparing a university's Azure-hosted ChatGPT deployment with commercial ChatGPT. There are no results, and the authors say so plainly. Judged as a proposal, it is competent and readable; judged as a contribution, it is a step toward one, not a step itself.\n\nWhat it does well: the compilation of LLMaaS customization categories (Table 1) is a useful organizing artifact, and the hypotheses are each traceable to prior work on trust cues, logos, warnings, and energy feedback. The survey instrument borrows validated measures where available and is explicit about which items are self-developed. For a work-in-progress report, it is honestly scoped and the writing is clear.\n\nThe soft spots are real. The stress-test note is on the mark: Table 2 lists substantive differences between the university system and ChatGPT that go well beyond user-salient cosmetic customizations—temperature fixed at 0, multiple models, image generation, GDPR-only EU processing, VPN-restricted access, token-percentage display, and training materials. The text says \"customizations regarding the data and model were kept to a minimum,\" which Table 2 contradicts. The planned observational comparison cannot cleanly attribute observed trust or usage differences to branding or interface design, because any of these features could drive effects. That is not a minor issue; it undermines the core contrast H1–H6 rely on. I do not see a statistical or design control for these dimensions in Section 3. The survey-items section is also thin—measure names and one example per construct, with the full instrument deferred to supplementary material that is not included. Fixing both before the field study would make the eventual results much more defensible. One small thing: reference [46] links to an example.com placeholder, which should be corrected.\n\nWho should read this: people designing or evaluating organizational LLM deployments, and HCI researchers working on trust cues. The paper earns serious referee time as a position or late-breaking work, but it would be a hard sell as a full paper without data. I would send it to review in a venue that explicitly welcomes extended abstract-and-plan formats, with the confound issue flagged for revision.","headline":"A clear, honest design prequel for a field study on how university-branded LLM customizations affect trust—worth engaging as a proposal, but the planned comparison currently cannot isolate the branding effect from several substantive system differences.","tokens_in":13099,"tokens_out":1455,"would_cite":false,"duration_ms":16494,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that user-visible LLM customizations—branding, interface, warnings, token display—act as trust cues that reshape trust, caution, privacy perception, and usage compared with commercial ChatGPT.","keywords":["LLM as a service","customization","trust cues","hallucinations","university AI","sustainable AI use","technology acceptance","field study"],"falsifier":"A randomized experiment in which otherwise identical LLM responses are presented to users under two conditions—one with university branding and warning text, one without—would settle the claim: if trust, verification behavior, and reported hallucination rates do not differ between conditions, the central cue argument is falsified. The planned non-random field comparison cannot distinguish branding from the other feature differences.","tokens_in":12198,"feed_emoji":"🎓","tokens_out":6725,"duration_ms":59548,"temperature":0.7,"pith_summary":"This paper is trying to establish that the visible, user-facing choices an organization makes when deploying an LLM-as-a-Service system—branding, layout, warnings, token displays—change how much people trust and how they use the system, even when the underlying model is unchanged. It proposes that these customizations act as trust cues that users read as signals of the system's capability and the organization's intentions. The authors back this with six testable hypotheses comparing a university's customized ChatGPT-based chatbot with commercial ChatGPT, and with a planned field-study survey. Since this is a late-breaking prequel, no empirical results are yet reported; the contribution is the argument, the hypotheses, and the measurement plan. If the argument holds, organizations can shape AI adoption, caution, privacy perception, and sustainability behavior through interface choices rather than model changes.","feed_headline":"University AI branding may shift trust and usage habits","feed_subtitle":"A planned field study compares a university-customized ChatGPT with the commercial version to test six trust and usage hypotheses.","key_machinery":"The load-bearing mechanism is the trust cue: any information element a user can use to make a trust assessment about an agent. The paper argues that user-salient LLMaaS customizations—university logo, corporate design, hallucination warnings, and token-percentage display—function as trust cues that map onto the process and purpose dimensions of trust in automation. The planned field study operationalizes this by asking the same participants to rate the customized chatbot and commercial ChatGPT on parallel items measuring trust, hallucination caution, privacy concern, and sustainability behavior, with organizational trust as a moderator.","core_discovery":"The central claim, on the paper's own terms, is that end users interpret customizations of LLMs as evidence about the system's capabilities, goals, and inner workings, even when those customizations, like corporate branding, do not affect functionality. From this, the authors derive six predictions: users will trust the customized university chatbot more than ChatGPT (H1); this effect is stronger when users trust the university itself (H2); users will behave less cautiously about hallucinations and will report encountering fewer of them (H3 and H4); users will feel greater privacy with the customized system (H5); and users will prompt more resource-efficiently and think more about sustainability when token usage is displayed (H6). The paper presents this as a research program rather than a result, positioning the present work as a design and hypothesis paper for a planned field study at a German university.","pith_inferences":["If branding alone moves trust, the same cue mechanism should transfer to other branded as-a-service AI deployments, such as corporate and government chatbots, though the paper does not test this.","The cleanest test of the core claim would be a randomized A/B comparison of two otherwise identical chatbot versions differing only in branding and warning placement; the planned field comparison includes additional differences, such as access rules, available models, output randomness, and token display, that could themselves affect trust and usage.","Hallucination caution can be measured behaviorally in usage logs, for example through clicks on verification or source links, rather than only by self-report; the current survey design relies on self-report."],"forward_implications":["Universities can raise initial trust and adoption of their AI services through branding and interface choices without improving the model itself.","Visible institutional branding may reduce users' critical checking, leading to overtrust and more uncorrected hallucinated content.","Displaying token usage as a percentage could nudge users toward more economical prompting, supporting sustainability goals.","Privacy perceptions can be raised by customization even when data handling is unchanged, which may mask real privacy tradeoffs.","The effect of customization on trust depends on pre-existing organizational trust, so institutions with weaker reputations may not benefit."],"supporting_citations":[{"why":"Supplies the definition of trust cues on which the whole argument rests.","marker":"[21]"},{"why":"Provides the taxonomy of LLMaaS customization options that separates user-salient from technical customizations.","marker":"[22]"},{"why":"Frames trust in automation along performance, process, and purpose dimensions, which the branding argument targets.","marker":"[34]"},{"why":"Shows chatbot trust is influenced by trust in the providing company, the basis for H2.","marker":"[44]"},{"why":"Provides evidence that familiar logos raise perceived credibility, the cue basis for branding effects.","marker":"[37]"},{"why":"Shows warnings help users detect hallucinations, supporting H3 and H4.","marker":"[43]"},{"why":"Supplies the organizational trust measure used to test the moderation in H2.","marker":"[12]"},{"why":"Supplies the trust-in-AI items used to measure the main dependent variable.","marker":"[63]"},{"why":"Shows transparent sustainability information changes user choices, supporting H6.","marker":"[49]"},{"why":"Shows real-time energy data promotes responsible consumption, supporting the token-visibility effect.","marker":"[52]"}],"fun_headline_variants":["University AI branding may sway trust and usage","How LLM branding alone could shape trust and behavior","Customized AI chatbots: trust via branding, not function","Branding AI: trust cues beyond the model","Campus LLM branding may alter trust and usage"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison assumes the university's chatbot and commercial ChatGPT differ only in the user-visible customizations under study; in fact they also differ in network access, available models, output randomness, and token display, any of which could independently influence trust and usage.","fun_headline_variants_meta":{"raw":{"variants":["University AI branding may sway trust and usage","How LLM branding alone could shape trust and behavior","Customized AI chatbots: trust via branding, not function","Branding AI: trust cues beyond the model","Campus LLM branding may alter trust and usage"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000801,"raw_usage":{"total_tokens":3497,"prompt_tokens":898,"completion_tokens":2599,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":2525}},"tokens_in":514,"tokens_out":2599,"duration_ms":18826,"temperature":1.0,"reasoning_tokens":2525,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:07:45.888237+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A randomized experiment in which otherwise identical LLM responses are presented to users under two conditions—one with university branding and warning text, one without—would settle the claim: if trust, verification behavior, and reported hallucination rates do not differ between conditions, the central cue argument is falsified. The planned non-random field comparison cannot distinguish branding from the other feature differences.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the definition of trust cues on which the whole argument rests."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the taxonomy of LLMaaS customization options that separates user-salient from technical customizations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides evidence that familiar logos raise perceived credibility, the cue basis for branding effects."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the organizational trust measure used to test the moderation in H2."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the trust-in-AI items used to measure the main dependent variable."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows transparent sustainability information changes user choices, supporting H6."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows real-time energy data promotes responsible consumption, supporting the token-visibility effect."}],"review_version":1}