{"id":"1df9f6e2-e4c9-4eca-9eea-042643b917e8","arxiv_id":"2506.14809","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A deployed LLM survey creation tool became the second most-used build method and produced higher activation and deployment rates than most traditional methods.","lead":"This paper reports on a large survey platform's LLM-powered survey creation tool, deployed in production and evaluated through the DeLone and McLean IS Success Model. It describes adoption, user satisfaction, and quality-control metrics, offering a case study for how generative AI can be integrated into a core research workflow.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Selection bias and missing denominators undermine the causal claim that Build with LLM drives higher activation and deployment.","rationale":"The reader's weakest assumption correctly identifies user comparability and self-selection as the pivotal issue. My analysis sharpens this by adding two concrete, missing evidentiary elements: the absence of any denominator or uncertainty information, and the absence of even a basic stratification that could separate tool effect from user context. These gaps directly threaten the causal reading of the system-use comparison, but they do not invalidate the paper's value as an experience report of a deployed system. The paper's other contributions—the hybrid evaluation framework, the TAM-based satisfaction analysis, and the post-deployment safeguards—stand independently and are described with reasonable detail. The information-quality section is admittedly incomplete, with human-evaluation scores omitted and citation errors, but that is not the load-bearing claim under review. Given the reader already returned CONDITIONAL, my concern reinforces that conditionality rather than moving the verdict to rejection. The appropriate action is to require the authors to supply the stratified comparison or temper the causal language in revision; hence the verdict remains CONDITIONAL, meaning the reader's original disposition is unchanged.","tokens_in":9780,"tokens_out":2071,"duration_ms":27413,"concrete_test":"Request the authors to provide the underlying per-survey counts and to re-estimate activation and deployment rates for Build with LLM versus the other non-copy methods stratified by at least two proxies for selection: (a) whether the user had created a survey on the platform before the study period, and (b) whether the survey topic falls into a simple one-off category (e.g., poll, quiz) versus a complex recurring category (e.g., employee engagement, longitudinal study). If the Build with LLM advantage disappears or reverses within the largest strata, the claim of downstream value is not robust to self-selection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the System Use section is that surveys created via Build with LLM achieve activation and deployment rates of 48.6% and 16.8%, and that among the four non-copy methods this indicates strong downstream value. This inference rests on an unstated comparability assumption: users who choose the LLM method are otherwise similar to users of the other methods. The paper provides no controls for user expertise, prior survey-creation experience, survey complexity, or task purpose, and it reports only aggregate rates with no sample sizes, confidence intervals, or significance tests. The authors themselves acknowledge that Copy from another survey has the highest rates because its users are conducting repeated longitudinal studies, which shows method choice is correlated with user context. By the same logic, users of Build with LLM may be early adopters, AI enthusiasts, or users with more ambitious survey projects, and the observed higher rates could reflect these pre-existing differences rather than any causal contribution of the tool. The absence of denominator information is especially important: if the Build with LLM group is small, a few projects with many responses could drive the deployment rate. As written, the comparison supports an association, not the causal 'helps more users move from initial survey creation to full survey deployment' statement.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports on a deployed LLM-based survey creation tool within a large commercial survey platform. Using the DeLone and McLean IS Success Model, it describes service quality and safeguards, analyzes 557 self-selected feedback responses for satisfaction and TAM constructs, compares activation and deployment rates across five survey creation methods over April 2024 to April 2025, and presents a hybrid human-automated evaluation of four LLM/prompt combinations. The central empirical claims are that surveys built with the LLM tool show the highest activation (48.6%) and deployment (16.8%) rates among the four non-copy methods, that promoters and detractors express distinct TAM-related themes, and that drift testing plus expert review supported selecting GPT4 with system prompt V2.","tokens_in":9942,"tokens_out":4300,"duration_ms":44682,"significance":"If the findings are valid, this is a useful real-world case study: it applies a mature IS framework to a generative AI tool, uses large production datasets (234,300 prompt-survey pairs for the acceptance classifier), proposes a reusable hybrid evaluation framework, and documents post-deployment safeguards. The paper's strength lies in its positioning as a deployed system evaluation rather than a lab study. However, the main quantitative claims rest on aggregate, self-reported metrics without external validation or statistical inference, which limits the strength of the conclusions until the missing validation is supplied.","major_comments":[{"comment":"The claim that Build with LLM 'demonstrates strong downstream value, helping more users move from initial survey creation to full survey deployment' is not warranted by the data presented. The comparison of aggregate activation and deployment rates across methods is observational, and the authors acknowledge for Copy from another survey that method choice is tied to user context (repeated longitudinal studies). The same selection-bias concern applies to Build with LLM: early adopters, AI-oriented users, or users with more ambitious projects may choose the LLM method, and the reported rates are not adjusted for user expertise, survey complexity, or task type. No denominators, confidence intervals, or significance tests are provided, so the reader cannot determine whether 48.6% and 16.8% are stable estimates or artifacts of small groups. Please report counts and uncertainty, and soften the causal language or add an appropriate quasi-experimental analysis.","section":"System Use, 'Usage Comparison of LLM System with Other Survey Creation Methods' (Figures 4-5)"},{"comment":"The TAM coding of open-ended feedback is performed by 'a second LLM,' but the manuscript provides no validation of this coding against human judges, no inter-coder reliability, and no details on the coding prompt or the number of comments in each segment. The percentages (62% PEOU, 55% PU, 48% BI for promoters; 41% usability, 34% UI, 29% editing, 25% transparency for detractors) are therefore unverified outputs of the coding model, not established measurements. The satisfaction analysis is also based on self-selected respondents (557 responses) with no response rate or comparison of respondents to non-respondents; this should be stated as a limitation and the TAM results should be labeled exploratory unless human validation is added.","section":"User Satisfaction, 'User Satisfaction' subsection"},{"comment":"The deployment decision to select GPT4 with system prompt V2 is justified by 'triangulat[ing] with human evaluation results,' but the human evaluation results are never reported. The paper lists the evaluation checklist (Table 4) and describes expert scoring, but no scores, number of rated surveys, number of raters, or inter-rater agreement appear anywhere. The PSI drift tests in Tables 6 and 7 establish only distributional differences between model/prompt versions, not which output is better. Without the human evaluation evidence, the central information-quality claim that GPT4+V2 is 'superior' is unsupported.","section":"Information Quality, 'Automated Evaluation of the LLM System' (Tables 4, 6, and 7)"},{"comment":"The binary classification analysis uses a LightGBM model on 234,300 prompt-survey pairs and draws conclusions about which features predict acceptance, but the manuscript reports no classification performance metrics such as accuracy, AUC, precision, recall, or calibration. If the model is not demonstrably predictive, the feature-importance claims (shorter prompts and more concise questions lead to acceptance) rest on an unvalidated instrument. Please report the model's evaluation metrics and, if appropriate, error bars on feature importances.","section":"Understanding User Behavior Through Binary Classification"}],"minor_comments":[{"comment":"The text says 'Figures 3 and 4 present the activation and deployment rates,' but the activation rate is Figure 4 and the deployment rate is Figure 5; please correct the cross-reference.","section":"System Use"},{"comment":"The text cites Papineni et al. (2018) for Self-BLEU, but the reference list entry is the original 2002 BLEU paper; please update the citation to the correct source for Self-BLEU.","section":"References"},{"comment":"The filtering steps exclude prompts with fewer than 200 or more than 500 characters and surveys with fewer than 5 or more than 12 questions; the rationale is given as distribution analyses, but the resulting evaluation set's representativeness of all user prompts should be discussed.","section":"Information Quality, 'Data Collection for LLM System Evaluation'"},{"comment":"Table 5's 'Aggregation' column lists 'count' for features such as any_special_character and score_flesch_kincaid; this is likely an artifact of the table formatting and should be clarified as the feature value or a summary statistic.","section":"Information Quality, 'Automated Evaluation of the LLM System' (Table 5)"},{"comment":"The paper would benefit from a data availability statement or a note that the production data cannot be shared for confidentiality reasons; currently no reproducibility information is provided.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an insider case study by the system's builders and the platform's research team. That is not disqualifying, but it increases the need for independent validation of the headline metrics; the current draft presents the system's success in fairly strong terms. If the authors can add validation and reframe claims, it could be a suitable contribution for the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this for the deployment data, not for the causal story. It is the first application of the IS Success Model to an LLM survey-creation tool, and it comes from a production system with real numbers: one year of usage, 234,300 prompt-survey pairs, 557 satisfaction responses, and a five-method comparison. The descriptive results—Build with LLM is the second most-used method and has the highest activation and deployment rates among the non-copy methods—are plausible and worth having on record.\n\nWhat the paper does well: the hybrid evaluation framework (human expert scoring plus PSI drift tests over model and prompt versions) is a practical, honest way to manage LLM updates; the safeguards section on prompt leakage and rate limiting is grounded in actual incidents; and the TAM-based qualitative coding of open-ended feedback, though not validated, is a reasonable exploratory step.\n\nThe soft spots are real. The comparison of activation and deployment rates is the centerpiece, but the authors treat it as evidence of the tool's causal contribution without controlling for who chooses the LLM method. The stress-test note is correct: the paper itself acknowledges that Copy from another survey has the highest rates because users are running longitudinal studies, which proves method choice is correlated with user context. By the same logic, Build with LLM users could be early adopters or more ambitious surveyors, and the higher rates could follow from that. No denominators, confidence intervals, or significance tests are reported. The causal wording in Section System Use should be downgraded to association.\n\nOther gaps: the human evaluation scores are promised but never reported; the LightGBM classifier is described but no accuracy or AUC is given; the TAM coding by an LLM is not checked against human coders; and there are citation inconsistencies (Papineni listed as 2018 in text, 2002 in references; the PSI reference is actually Taplin and Hunt's population accuracy index, a different measure). These are fixable but need addressing.\n\nWho this is for: IS/HCI researchers and practitioners who want realistic adoption metrics from a deployed genAI feature. It deserves a serious referee, but only with revision: report the missing statistics, supply denominators for the rates, and soften the causal claims. I'd accept it for the record as an experience report; I would not cite it for anything stronger than an existence proof of adoption.","headline":"A credible industry experience report with real deployment numbers, but the central causal claim about downstream value is undermined by unmeasured selection effects and missing statistics.","tokens_in":10487,"tokens_out":3037,"would_cite":false,"duration_ms":31997,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A deployed LLM survey builder, evaluated through the DeLone and McLean IS Success Model, produces surveys that get shared and collect responses at the highest rate of any non-copy creation method.","keywords":["LLM survey generation","IS Success Model","DeLone and McLean","Technology Acceptance Model","generative AI evaluation","Population Stability Index","deployment safeguards","human-AI collaboration"],"falsifier":"A matched or randomized comparison in which users with similar survey topics, platform experience, and prompt characteristics are assigned to the LLM builder or to a manual or template method; if the LLM's activation and deployment advantage disappears or reverses, the paper's central attribution is falsified. A simpler observational check would compute activation and deployment rates separately for the 37% of users who use sample prompts versus the 63% who write their own prompts; if the advantage is concentrated in one group, the tool-level claim weakens.","tokens_in":9543,"feed_emoji":"📊","tokens_out":4785,"duration_ms":46901,"temperature":0.7,"pith_summary":"The paper reports on a large survey platform's LLM-powered survey builder, released in late 2023, and asks whether it delivers real value once deployed. Using the DeLone and McLean IS Success Model, it assembles evidence on service quality, user satisfaction, system use, and information quality. The central finding is that surveys built with the LLM are shared at a 48.6% activation rate and reach five or more responses at a 16.8% deployment rate, the highest among all methods except copying an existing survey. The paper argues this shows the tool moves non-expert users from creation into genuine data collection, and that the hybrid automated-plus-human evaluation framework lets the team choose safe updates. A sympathetic reader would take the contribution as a real-world template for evaluating generative AI systems inside an established IS theory.","feed_headline":"LLM-built surveys beat four methods at reaching live use","feed_subtitle":"Platform data: AI-generated surveys hit 48.6% activation and 16.8% deployment, second only to copying an old survey.","key_machinery":"The load-bearing object is the DeLone and McLean IS Success Model, applied here for the first time to an LLM survey generator; it supplies the four evaluation dimensions (service quality, user satisfaction, system use, and information quality) that organize the study. Within that frame, the paper's main quantitative instruments are adoption funnel metrics (activation and deployment), the Technology Acceptance Model used to code open-ended feedback into Perceived Ease of Use, Perceived Usefulness, and Behavioral Intention, and a hybrid information-quality framework in which expert human raters score generated surveys while automated Population Stability Index tests on metadata features flag distributional drift between model or prompt versions.","core_discovery":"On its own terms, the paper's discovery is that the LLM tool behaves like a successful information system, not just a novel feature: among the four non-copy creation methods, Build with LLM has the highest activation rate (48.6%) and deployment rate (16.8%) over April 2024 to April 2025, making it the second most-used method overall despite being the newest. Copy from another survey scores higher only because those users run repeated longitudinal studies with established instruments. The paper also finds moderate user satisfaction (3.18 out of 5), with promoters citing ease and usefulness and detractors citing reliability and UI friction, and it shows that switching from GPT-3.5 with prompt V1 to GPT-4 with prompt V2 changed survey-generation behavior enough to require drift testing before production adoption.","pith_inferences":["An implication the authors leave implicit is that the 48.6% and 16.8% advantage is correlational: users choose their creation method, and the comparison includes no control for topic complexity, platform tenure, or prior survey experience, so part of the advantage may be self-selection rather than a tool effect.","A testable extension would randomize or match users across creation methods; if the activation gap persists under assignment, the causal claim would be much stronger.","Because the paper finds that user profile features were not predictive of acceptance, a promising next step would be to test whether prompt wording alone, rather than who writes it, drives downstream success; this could be validated by comparing sample prompts against user-written prompts of similar length."],"forward_implications":["If the IS Success Model holds, the LLM tool's downstream value is evidence that generative survey creation can move users further along the data-collection pipeline than manual or template methods, aside from copying an existing survey.","Because the production deployment decision rested on the hybrid evaluation, the framework can be reused to gate future model and prompt updates before release.","The prompt-use analysis implies that the five sample prompts cover the dominant user needs, and the emergence of employee engagement and event registration categories suggests that adding those sample prompts would increase adoption.","The classification result that short prompts and concise questions predict survey acceptance gives a concrete, testable design rule for the prompt-engineering layer."],"supporting_citations":[{"why":"Supplies the IS Success Model that organizes the paper's four evaluation dimensions and frames the tool's value.","marker":"DeLone and McLean (2003)"},{"why":"Supplies the Technology Acceptance Model constructs used to code open-ended user feedback into PEOU, PU, and BI.","marker":"Davis (1989)"},{"why":"Provides LightGBM, the classifier used to predict whether users accept a generated survey.","marker":"Ke et al. (2017)"},{"why":"Supplies the Population Stability Index used to detect distributional drift between LLM and prompt versions.","marker":"Taplin and Hunt (2019)"},{"why":"Supplies the Flesch-Kincaid readability score used as a metadata feature in the automated evaluation.","marker":"Kincaid et al. (1975)"},{"why":"Previous application of the IS Success Model to ChatGPT, which this paper extends to the survey-creation domain.","marker":"Marjanovic et al. (2024)"},{"why":"Introduces SurveyX, an alternative LLM survey tool that the paper contrasts with its own deployed and deployment-grounded system.","marker":"Liang et al. (2025)"}],"fun_headline_variants":["LLM surveys top non-copy methods in live use: 48.6%","First IS Success Model applied to AI survey creator","AI survey tool: 48.6% activation, tops new methods","Second-most-used survey method: LLM tool after copy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Users who choose the LLM method are comparable to users who choose other methods, so the higher activation and deployment rates can be credited to the tool rather than to who picks it or what survey they are building.","fun_headline_variants_meta":{"raw":{"variants":["LLM surveys top non-copy methods in live use: 48.6%","First IS Success Model applied to AI survey creator","AI survey tool: 48.6% activation, tops new methods","Second-most-used survey method: LLM tool after copy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001832,"raw_usage":{"total_tokens":7166,"prompt_tokens":870,"completion_tokens":6296,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":6222}},"tokens_in":486,"tokens_out":6296,"duration_ms":46145,"temperature":1.0,"reasoning_tokens":6222,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:02:26.635148+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A matched or randomized comparison in which users with similar survey topics, platform experience, and prompt characteristics are assigned to the LLM builder or to a manual or template method; if the LLM's activation and deployment advantage disappears or reverses, the paper's central attribution is falsified. A simpler observational check would compute activation and deployment rates separately for the 37% of users who use sample prompts versus the 63% who write their own prompts; if the advantage is concentrated in one group, the tool-level claim weakens.","supporting_citations":[],"review_version":1}