{"id":"244aa172-efd5-4353-98fe-5da251f6bb13","arxiv_id":"2608.00357","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A survey platform that uses LLMs to cluster open-ended responses live and lets respondents rate, rank, and explain those clusters produced positive user reactions in two small field studies, though the 'richer than traditional surveys' claim remains unproven.","lead":"The paper introduces Dynamic Surveys, an LLM-powered survey platform that clusters open-ended answers in real time and asks respondents to rate, rank, and reflect on those clusters. In two small field studies, participants reported feeling more engaged and said they shared deeper thoughts, but the comparison to traditional survey tools relies on self-report rather than a controlled test.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central comparative claim is untested: no control/baseline arm, and the only quantitative evidence is self-selected post-study self-report, so the observed ratings cannot support 'richer/deeper insights' vs traditional surveys.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the central comparative claim rests on self-selected post-study self-reports with no baseline condition. My reading of the full manuscript confirms this. Section 4.2 describes the post-study survey, Section 5 reports the results, and Section 7 explicitly concedes acquiescence bias and conflated item wording. The absence of any control arm means that even perfectly accurate self-reports cannot support the comparative proposition in the abstract. I also considered whether another concern—cluster accuracy without inter-rater reliability, or the unreported correlation in Section 5.1.3—might be more central. These are real but secondary; they affect internal quality metrics, not the headline comparative claim. The system description is coherent, prompts are included, and the limitations section is candid, so the paper has genuine merit as a system and feasibility study. However, the comparative benefits are not established by the current evidence. A controlled experiment with blind-coded depth ratings and behavioral engagement measures would settle the concern. Since this matches the reader's conditional verdict, no adjustment is needed.","tokens_in":20855,"tokens_out":2841,"duration_ms":31345,"concrete_test":"Run a preregistered between-subjects experiment with the same research question across three arms: Dynamic Surveys, a traditional open-ended survey, and a closed-ended Likert survey, recruited from the same population (e.g., n≈50 per arm). Blind to condition, independent coders rate response depth using a pre-specified rubric with inter-rater reliability (e.g., Krippendorff's alpha ≥ .80); measure behavioral engagement (completion rate, time-on-task, optional elaboration length/word count) and post-survey perceptions. Pre-specify that the comparative claim is supported only if Dynamic Surveys significantly outperforms both baselines on blind-coded depth and at least one behavioral engagement metric, not just on self-report. This directly tests whether the observed effects are attributable to the platform rather than to novelty or demand characteristics.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (abstract; Sections 5.3-5.4) is explicitly comparative: Dynamic Surveys provide 'richer and deeper insights' and 'increase engagement and foster a sense of community' relative to traditional survey tools. The evaluation does not include any baseline or control condition. All 93 participants experienced only Dynamic Surveys; the quantitative evidence for depth, engagement, and community comes from a post-study survey (Section 4.2) completed by 44 of 93 participants (47%), self-selected, using Likert agreement statements with no anchor to a specific prior survey experience (e.g., 'I shared thoughts and opinions in a greater depth than other feedback surveys'). Section 5 reports these as percentages without a comparison arm. The authors themselves acknowledge acquiescence bias and conflated dimensions in Section 7. Because the outcome measures are subjective self-reports from a subset of participants who just used a novel system, the observed ratings are equally consistent with novelty effects, demand characteristics, or social desirability as with genuine superiority. This is the load-bearing flaw: even if every reported rating accurately reflects participant feeling, the data cannot establish the comparative claim because there is no counterfactual. The stakeholder interviews (n=4) are suggestive but not a substitute for a controlled comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Dynamic Surveys, an LLM-powered survey platform that clusters open-ended responses in real time, generates personalized follow-up questions, and asks respondents to rate, rank, and explain their ranking of emergent themes. The authors claim that this design yields richer and deeper insights than traditional survey tools and increases engagement and community feeling. They report two field studies with 93 student respondents (plus four stakeholder interviewees), a voluntary post-study survey completed by 44 respondents, and manual cluster-accuracy coding. The system design is described in detail, including prompts in an appendix. The central comparative claims, however, rest on an evaluation that lacks a baseline condition and uses self-selected self-report data, which cannot support the strength of the conclusions drawn.","tokens_in":21155,"tokens_out":4461,"duration_ms":40860,"significance":"If the claims were well supported, the platform would be a meaningful contribution to CSCW/HCI: it offers a lightweight, scalable blend of qualitative and quantitative data collection with a participatory layer, and the design space is timely. The paper provides a useful system description, transparent prompts, and thoughtful design implications, and it explicitly acknowledges some limitations. However, the current evidence is exploratory at best. The lack of a control or baseline condition, the circularity in the follow-up evaluation, and the absence of inter-rater reliability for cluster-accuracy coding mean that the abstract's comparative statements are not warranted. The contribution is best framed as a design proposal with preliminary feasibility data, not as a demonstrated improvement over existing survey tools.","major_comments":[{"comment":"The paper's headline claim is explicitly comparative: 'richer and deeper insights compared with traditional survey tools' and 'increase engagement and foster a sense of community.' Yet no baseline or control condition exists. All 93 participants experienced only Dynamic Surveys. The post-study survey (Section 4.2) was completed by 44 of 93 (47%) self-selected respondents and uses Likert agreement items with implicit comparisons (e.g., 'I shared thoughts and opinions in a greater depth than other feedback surveys') without anchoring to a specific prior survey experience. Section 5.3 and 5.4 report these ratings as evidence, but they are equally consistent with novelty effects, demand characteristics, or acquiescence bias. The authors acknowledge acquiescence in Section 7 but not the absence of a counterfactual. This is a load-bearing flaw for the central claim and must be fixed by reframi","section":"Abstract; Sections 4.2, 5.3, 5.4"},{"comment":"The evaluation of follow-up responses reuses the same GPT clustering prompt that runs in the platform. That prompt explicitly instructs the model to create new themes not covered by existing clusters (guideline 3 and step 2 in Appendix A.1). Finding 'additional themes' in follow-up responses is therefore a consequence of the prompt design, not independent evidence that follow-up questions provide new insight. This check can demonstrate internal consistency of the LLM clustering, but it cannot validate the claim that follow-up responses enrich the data beyond what the original open-ended responses would have yielded. The depth contribution in Section 5.1.4 is thus circular.","section":"Section 4.4 and Appendix A.1"},{"comment":"Cluster accuracy is manually coded by members of the research team with no inter-rater reliability statistic reported. The text says 'multiple members of the research team, who independently coded responses and then discussed differences to reach consensus,' but no Kappa or agreement coefficient is given. Without this, the accuracy scores (83.8 and 92.7) cannot be distinguished from subjective judgment, especially given the authors' stake in the system. This weakens the quantitative evidence for cluster quality and should be reported or at least explicitly acknowledged as a limitation.","section":"Section 5.1.2 and Tables 3-4"},{"comment":"The claimed 'positive correlation' between cluster ranking scores and agreement proportions is not supported by any statistical test or correlation coefficient. The text notes alignment in the top half of clusters but not the lower half and attributes this to later cluster formation, but no timing data or significance test is presented. The assertion is therefore purely descriptive, and the proposed explanation is speculative. If this relationship is used to argue for the validity of the ranking scores, it should be tested (e.g., Spearman's rho) and the timing hypothesis examined with actual cluster-creation timestamps.","section":"Section 5.1.3 and Figure 4"},{"comment":"The limitations section acknowledges the single-topic focus, narrow demographics, limited stakeholder sample, and bias in scale wording, but it omits the most fundamental limitation for the abstract's claims: the absence of any baseline condition. Since the paper claims superiority over traditional survey tools, the lack of a control arm must be stated explicitly, and the conclusions must be tempered accordingly. The current limitations section gives readers the impression that the main threats are minor when the central comparative claim is untested.","section":"Section 7 (Limitations)"}],"minor_comments":[{"comment":"The total participant count is inconsistent: the abstract and Section 1 say 93 participants, while Section 4.1 says 97 including the four stakeholders. Please clarify whether the 93 refers to respondents only and 97 to all study participants.","section":"Section 1 / Abstract"},{"comment":"The word 'protostudy' appears without a hyphen; should be 'proto-study' for readability.","section":"Section 3.2"},{"comment":"The body text says 'agreement proportions' but Figure 4 caption defines 'agree' as including 'somewhat agree' and 'strongly agree.' The wording in the text should match the figure caption and clearly define the response categories used.","section":"Section 5.1.3"},{"comment":"The coding scheme for cluster accuracy (0, 0.5, 1 points) should be accompanied by a more explicit description of how 'maximum possible points' was computed, so that the accuracy percentage is reproducible.","section":"Section 5.1.2"},{"comment":"Consider adding the number of coders and an inter-rater reliability statistic (e.g., Cohen's kappa) directly in the table caption, rather than only mentioning consensus in the text.","section":"Tables 3-4"},{"comment":"There is a duplicated sentence: 'Recent research has begun exploring how surveys can support greater interactivity.' appears twice in consecutive paragraphs. Remove one occurrence.","section":"Section 2.1"},{"comment":"The phrase 'may lead to more deeper insights' should be corrected to 'deeper insights.'","section":"Section 6.1"}],"recommendation":"major_revision","confidential_remarks":"This paper is on the borderline between major revision and reject. The core system design and qualitative observations are potentially useful to the CSCW community, but the central comparative claims are unsupported by the current evaluation design. The lack of a baseline condition cannot be fixed retrospectively; the authors must substantially reframe the contribution as a design exploration with preliminary feasibility evidence. If they are unwilling to temper the abstract and conclusions in this way, I would recommend rejection. The circularity in the follow-up evaluation and the missing inter-rater reliability further weaken the evidentiary basis, but these could be addressed in a revised version if the claims are scaled back."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real contribution here is the design: real-time LLM clustering of open-ended responses, followed by rating, ranking, and reflective comparison inside the survey flow. I have not seen that combination in the cited systems, and the authors did the field work to show it can run in practice. The appendix prompts, the Borda score with its confidence-bound modification, and the explicit boundary that this is for early-stage exploration, not interpretive qualitative paradigms, are all to their credit. The paper is honestly written and the system description is clear enough to reproduce.\n\nThe soft spot is exactly where the reader put it. The abstract and Sections 5.3-5.4 claim richer, deeper insights and greater engagement compared with traditional survey tools. That comparison is never tested. There is no baseline arm; all 93 participants experienced only Dynamic Surveys. The quantitative evidence is a post-study survey completed by 44 self-selected respondents, using agreement statements with no concrete anchor to a prior survey experience. The authors acknowledge acquiescence bias in Section 7, but acknowledgement does not carry the load of the claim. The four stakeholder interviews are suggestive, not a counterfactual.\n\nTwo further soft spots, in proportion. The Section 4.4 check that follow-up responses add new themes reuses the exact GPT clustering prompt from the platform, so it can only demonstrate internal consistency, not external value. The cluster-accuracy coding reports no inter-rater reliability, only team consensus. No code or data are released, which limits the ability of others to build on the platform directly.\n\nNone of this sinks the design contribution. The paper is a solid systems-and-field-study report for HCI/CSCW readers working on LLM-based data collection. It deserves a serious referee, not a desk rejection. The path is straightforward: re-scope the headline claims to perceived experience and feasibility, or add a controlled comparison against a conventional survey, and report agreement metrics and release artifacts. I would engage with it as a reviewer, but I would not cite the comparative benefits as established until that revision happens.","headline":"A genuinely novel integrated survey design, undermined by a comparative claim the evidence was never built to test.","tokens_in":21633,"tokens_out":2430,"would_cite":false,"duration_ms":27020,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Dynamic Surveys uses LLMs to cluster open-ended answers as they arrive and lets respondents rank and reflect on the themes.","keywords":["dynamic surveys","LLM clustering","open-ended responses","qualitative-quantitative blending","respondent engagement","participatory surveys","thematic analysis","exploratory research"],"falsifier":"Run a preregistered randomized experiment in which respondents are assigned either to Dynamic Surveys or to a conventional open-ended-plus-Likert survey on the same topic; if independent coders blind to condition rate the responses for depth and specificity and find no advantage for Dynamic Surveys, or if completion rates are not higher, the paper's central claim is refuted.","tokens_in":20767,"feed_emoji":"📊","tokens_out":7441,"duration_ms":65584,"temperature":0.7,"pith_summary":"This paper proposes a survey method that moves the work of interpreting free-text responses into the survey itself. Respondents answer one open-ended question, receive a personalized follow-up, and then rate and rank the thematic clusters that an LLM has built from everyone's answers so far. The authors argue this blends the depth of open-ended responses with the structure of closed-ended questions, and that letting respondents see and react to emerging themes increases engagement and a sense of community. Evidence comes from two field studies with 93 students plus interviews with four stakeholders; the paper reports that clusters were largely accurate, follow-up answers surfaced new themes, and participants said they thought harder and felt more connected. If the claim holds, it gives non-expert researchers a lightweight way to surface and validate themes during collection rather than after it.","feed_headline":"LLMs cluster survey answers live, and respondents rank the themes","feed_subtitle":"A new platform blends open-ended depth with structured ranking; two field studies with 93 students report deeper engagement.","key_machinery":"The load-bearing mechanism is the real-time clustering loop: a structured prompt assigns each new response to existing themes and generates new ones only for core ideas not yet covered, with explicit instructions to avoid overlap and fragmentation. This loop feeds two other components—a personalized follow-up question generated from the respondent's own answer, and a ranking step that uses a positional voting rule (points assigned by rank order) with a lower-confidence-bound adjustment so that clusters seen by few respondents are not overvalued. The final component is the public report page, which aggregates clusters into a ranked list with opinion distributions, summaries, and minority-voic","core_discovery":"The central claim is that LLM-powered clustering can be a data-collection instrument rather than a post-hoc analysis tool. As each response arrives, the platform's LLM first places it into existing thematic clusters and then creates new clusters for ideas those themes do not yet cover. Those live clusters are shown back to respondents, who rate each on a five-point scale, rank them by value to the group, and write explanations when their ranking differs from the aggregate. The report page then presents ranked clusters with summaries, opinion distributions, original responses, and minority explanations. Across the two studies, the generated clusters received an average manual accuracy score o","pith_inferences":["A fair test of the 'richer insights' claim would require independent coders, blind to condition, to rate response depth against a conventional survey; the paper's self-report evidence cannot rule out novelty effects.","The clustering-plus-ranking loop effectively crowd-sources a lightweight thematic analysis, so the same pattern could be adapted to other structured elicitation tasks, such as requirements gathering or participatory budgeting, where participants react to emergent categories.","The sense of community the paper attributes to seeing others' responses is a testable design property: one could measure whether it persists when clusters are hidden until after submission, when respondents are fully anonymous to one another, or when the respondent pool is large enough to make clusters unstable.","The confidence-bound adjustment for clusters seen by few respondents suggests a practical stopping rule: keep the survey open until every cluster has received a target number of ratings, which would reduce the bias the authors observe in lower-ranked clusters."],"forward_implications":["Survey creators without qualitative training could get structured thematic reports from free-text responses without hiring coders or doing post-hoc analysis.","Respondents produce deeper input: personalized follow-up questions and the request to explain ranking differences elicit concrete examples, reasoning, and suggestions that original answers omit.","The public report turns a one-way survey into a lightweight, asynchronous exchange, letting respondents see where their views converge with and diverge from the group.","The method is positioned for exploratory and directional insight, not for statistical generalization or confirmatory theory building, making it a complement to interviews rather than a replacement.","The findings point to design constraints for future LLM survey tools: balancing question depth against respondent burden, controlling cluster proliferation, and giving creators real-time quality indicators."],"fun_headline_variants":["Live LLM clustering turns open-ended answers into ranked themes","Respondents rank LLM-made clusters and explain disagreements","Surveys that cluster answers live, users rank and react","Dynamic Surveys: real-time clusters, user rankings, richer insight","LLM blends open-ended depth with quantitative ranking in surveys"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The comparison to traditional survey tools rests on self-reported ratings from 44 of the 93 respondents plus four stakeholder interviews, so if those reports reflect politeness, novelty, or leading questions rather than actual gains in insight and engagement, the central claim loses its empirical support.","fun_headline_variants_meta":{"raw":{"variants":["Live LLM clustering turns open-ended answers into ranked themes","Respondents rank LLM-made clusters and explain disagreements","Surveys that cluster answers live, users rank and react","Dynamic Surveys: real-time clusters, user rankings, richer insight","LLM blends open-ended depth with quantitative ranking in surveys"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000142,"raw_usage":{"total_tokens":1039,"prompt_tokens":815,"completion_tokens":224,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":142}},"tokens_in":559,"tokens_out":224,"duration_ms":3047,"temperature":1.0,"reasoning_tokens":142,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T00:37:13.346638+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a preregistered randomized experiment in which respondents are assigned either to Dynamic Surveys or to a conventional open-ended-plus-Likert survey on the same topic; if independent coders blind to condition rate the responses for depth and specificity and find no advantage for Dynamic Surveys, or if completion rates are not higher, the paper's central claim is refuted.","supporting_citations":[],"review_version":1}