Pith. sign in

REVIEW 5 major objections 7 minor 86 references

Dynamic Surveys uses LLMs to cluster open-ended answers as they arrive and lets respondents rank and reflect on the themes.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-08-04 00:37 UTC pith:TMKJ5L3J

load-bearing objection A genuinely novel integrated survey design, undermined by a comparative claim the evidence was never built to test. the 5 major comments →

arxiv 2608.00357 v1 pith:TMKJ5L3J submitted 2026-08-01 cs.HC

Dynamic Surveys: Using LLMs to Blend Qualitative Depth,Quantitative Structure, and Collaborative Interaction

classification cs.HC
keywords dynamic surveysLLM clusteringopen-ended responsesqualitative-quantitative blendingrespondent engagementparticipatory surveysthematic analysisexploratory research
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper proposes a survey method that moves the work of interpreting free-text responses into the survey itself. Respondents answer one open-ended question, receive a personalized follow-up, and then rate and rank the thematic clusters that an LLM has built from everyone's answers so far. The authors argue this blends the depth of open-ended responses with the structure of closed-ended questions, and that letting respondents see and react to emerging themes increases engagement and a sense of community. Evidence comes from two field studies with 93 students plus interviews with four stakeholders; the paper reports that clusters were largely accurate, follow-up answers surfaced new themes, and participants said they thought harder and felt more connected. If the claim holds, it gives non-expert researchers a lightweight way to surface and validate themes during collection rather than after it.

Core claim

The central claim is that LLM-powered clustering can be a data-collection instrument rather than a post-hoc analysis tool. As each response arrives, the platform's LLM first places it into existing thematic clusters and then creates new clusters for ideas those themes do not yet cover. Those live clusters are shown back to respondents, who rate each on a five-point scale, rank them by value to the group, and write explanations when their ranking differs from the aggregate. The report page then presents ranked clusters with summaries, opinion distributions, original responses, and minority explanations. Across the two studies, the generated clusters received an average manual accuracy score o

What carries the argument

The load-bearing mechanism is the real-time clustering loop: a structured prompt assigns each new response to existing themes and generates new ones only for core ideas not yet covered, with explicit instructions to avoid overlap and fragmentation. This loop feeds two other components—a personalized follow-up question generated from the respondent's own answer, and a ranking step that uses a positional voting rule (points assigned by rank order) with a lower-confidence-bound adjustment so that clusters seen by few respondents are not overvalued. The final component is the public report page, which aggregates clusters into a ranked list with opinion distributions, summaries, and minority-voic

Load-bearing premise

The comparison to traditional survey tools rests on self-reported ratings from 44 of the 93 respondents plus four stakeholder interviews, so if those reports reflect politeness, novelty, or leading questions rather than actual gains in insight and engagement, the central claim loses its empirical support.

What would settle it

Run a preregistered randomized experiment in which respondents are assigned either to Dynamic Surveys or to a conventional open-ended-plus-Likert survey on the same topic; if independent coders blind to condition rate the responses for depth and specificity and find no advantage for Dynamic Surveys, or if completion rates are not higher, the paper's central claim is refuted.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Survey creators without qualitative training could get structured thematic reports from free-text responses without hiring coders or doing post-hoc analysis.
  • Respondents produce deeper input: personalized follow-up questions and the request to explain ranking differences elicit concrete examples, reasoning, and suggestions that original answers omit.
  • The public report turns a one-way survey into a lightweight, asynchronous exchange, letting respondents see where their views converge with and diverge from the group.
  • The method is positioned for exploratory and directional insight, not for statistical generalization or confirmatory theory building, making it a complement to interviews rather than a replacement.
  • The findings point to design constraints for future LLM survey tools: balancing question depth against respondent burden, controlling cluster proliferation, and giving creators real-time quality indicators.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A fair test of the 'richer insights' claim would require independent coders, blind to condition, to rate response depth against a conventional survey; the paper's self-report evidence cannot rule out novelty effects.
  • The clustering-plus-ranking loop effectively crowd-sources a lightweight thematic analysis, so the same pattern could be adapted to other structured elicitation tasks, such as requirements gathering or participatory budgeting, where participants react to emergent categories.
  • The sense of community the paper attributes to seeing others' responses is a testable design property: one could measure whether it persists when clusters are hidden until after submission, when respondents are fully anonymous to one another, or when the respondent pool is large enough to make clusters unstable.
  • The confidence-bound adjustment for clusters seen by few respondents suggests a practical stopping rule: keep the survey open until every cluster has received a target number of ratings, which would reduce the bias the authors observe in lower-ranked clusters.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 7 minor

Summary. The paper introduces Dynamic Surveys, an LLM-powered survey platform that clusters open-ended responses in real time, generates personalized follow-up questions, and asks respondents to rate, rank, and explain their ranking of emergent themes. The authors claim that this design yields richer and deeper insights than traditional survey tools and increases engagement and community feeling. They report two field studies with 93 student respondents (plus four stakeholder interviewees), a voluntary post-study survey completed by 44 respondents, and manual cluster-accuracy coding. The system design is described in detail, including prompts in an appendix. The central comparative claims, however, rest on an evaluation that lacks a baseline condition and uses self-selected self-report data, which cannot support the strength of the conclusions drawn.

Significance. If the claims were well supported, the platform would be a meaningful contribution to CSCW/HCI: it offers a lightweight, scalable blend of qualitative and quantitative data collection with a participatory layer, and the design space is timely. The paper provides a useful system description, transparent prompts, and thoughtful design implications, and it explicitly acknowledges some limitations. However, the current evidence is exploratory at best. The lack of a control or baseline condition, the circularity in the follow-up evaluation, and the absence of inter-rater reliability for cluster-accuracy coding mean that the abstract's comparative statements are not warranted. The contribution is best framed as a design proposal with preliminary feasibility data, not as a demonstrated improvement over existing survey tools.

major comments (5)
  1. [Abstract; Sections 4.2, 5.3, 5.4] The paper's headline claim is explicitly comparative: 'richer and deeper insights compared with traditional survey tools' and 'increase engagement and foster a sense of community.' Yet no baseline or control condition exists. All 93 participants experienced only Dynamic Surveys. The post-study survey (Section 4.2) was completed by 44 of 93 (47%) self-selected respondents and uses Likert agreement items with implicit comparisons (e.g., 'I shared thoughts and opinions in a greater depth than other feedback surveys') without anchoring to a specific prior survey experience. Section 5.3 and 5.4 report these ratings as evidence, but they are equally consistent with novelty effects, demand characteristics, or acquiescence bias. The authors acknowledge acquiescence in Section 7 but not the absence of a counterfactual. This is a load-bearing flaw for the central claim and must be fixed by reframi
  2. [Section 4.4 and Appendix A.1] The evaluation of follow-up responses reuses the same GPT clustering prompt that runs in the platform. That prompt explicitly instructs the model to create new themes not covered by existing clusters (guideline 3 and step 2 in Appendix A.1). Finding 'additional themes' in follow-up responses is therefore a consequence of the prompt design, not independent evidence that follow-up questions provide new insight. This check can demonstrate internal consistency of the LLM clustering, but it cannot validate the claim that follow-up responses enrich the data beyond what the original open-ended responses would have yielded. The depth contribution in Section 5.1.4 is thus circular.
  3. [Section 5.1.2 and Tables 3-4] Cluster accuracy is manually coded by members of the research team with no inter-rater reliability statistic reported. The text says 'multiple members of the research team, who independently coded responses and then discussed differences to reach consensus,' but no Kappa or agreement coefficient is given. Without this, the accuracy scores (83.8 and 92.7) cannot be distinguished from subjective judgment, especially given the authors' stake in the system. This weakens the quantitative evidence for cluster quality and should be reported or at least explicitly acknowledged as a limitation.
  4. [Section 5.1.3 and Figure 4] The claimed 'positive correlation' between cluster ranking scores and agreement proportions is not supported by any statistical test or correlation coefficient. The text notes alignment in the top half of clusters but not the lower half and attributes this to later cluster formation, but no timing data or significance test is presented. The assertion is therefore purely descriptive, and the proposed explanation is speculative. If this relationship is used to argue for the validity of the ranking scores, it should be tested (e.g., Spearman's rho) and the timing hypothesis examined with actual cluster-creation timestamps.
  5. [Section 7 (Limitations)] The limitations section acknowledges the single-topic focus, narrow demographics, limited stakeholder sample, and bias in scale wording, but it omits the most fundamental limitation for the abstract's claims: the absence of any baseline condition. Since the paper claims superiority over traditional survey tools, the lack of a control arm must be stated explicitly, and the conclusions must be tempered accordingly. The current limitations section gives readers the impression that the main threats are minor when the central comparative claim is untested.
minor comments (7)
  1. [Section 1 / Abstract] The total participant count is inconsistent: the abstract and Section 1 say 93 participants, while Section 4.1 says 97 including the four stakeholders. Please clarify whether the 93 refers to respondents only and 97 to all study participants.
  2. [Section 3.2] The word 'protostudy' appears without a hyphen; should be 'proto-study' for readability.
  3. [Section 5.1.3] The body text says 'agreement proportions' but Figure 4 caption defines 'agree' as including 'somewhat agree' and 'strongly agree.' The wording in the text should match the figure caption and clearly define the response categories used.
  4. [Section 5.1.2] The coding scheme for cluster accuracy (0, 0.5, 1 points) should be accompanied by a more explicit description of how 'maximum possible points' was computed, so that the accuracy percentage is reproducible.
  5. [Tables 3-4] Consider adding the number of coders and an inter-rater reliability statistic (e.g., Cohen's kappa) directly in the table caption, rather than only mentioning consensus in the text.
  6. [Section 2.1] There is a duplicated sentence: 'Recent research has begun exploring how surveys can support greater interactivity.' appears twice in consecutive paragraphs. Remove one occurrence.
  7. [Section 6.1] The phrase 'may lead to more deeper insights' should be corrected to 'deeper insights.'

Circularity Check

1 steps flagged

Same-prompt re-clustering makes the 'new themes from follow-ups' finding a by-construction result; other claims rest on unblinded self-report, not circular.

specific steps
  1. self definitional [Section 5.1.4; Appendix A.1 (clustering prompt)]
    "To understand whether responses to follow-up questions could provide new insights into the original questions, we applied the same clustering method to the follow-up responses in both surveys. We found that these follow-up responses revealed additional themes ... If you believe that the current themes do not properly capture ALL core ideas of the response, create a new general theme or multiple new themes that fully capture the ideas in the response."

    The clustering prompt used in the platform explicitly instructs the model to create new themes whenever an unrepresented core idea appears. Applying that same prompt to follow-up responses therefore guarantees that any novel content will be emitted as a new theme; finding 'additional themes' is an execution of the prompt's built-in rule, not an independent test that follow-up adds depth. The result cannot validate the platform's value beyond internal consistency, yet it is reported as evidence that follow-up questions elicit deeper insights.

full rationale

The paper's central comparative claims (richer/deeper insights, engagement, community) are supported primarily by post-study self-reports and stakeholder interviews with no control arm; this is an empirical design limitation, not a circular derivation. The only load-bearing step that reduces by construction is the re-use of the platform's own clustering prompt (which mandates new-theme generation) to conclude that follow-up responses surface additional themes. Other quantitative analyses (cluster accuracy, ranking-agreement correlation) are descriptive and manually coded, not fitted predictions. No self-citation chain or imported uniqueness theorem is present. Overall, a partial circularity score of 4 reflects the by-construction sub-check while the main claim has independent (though weak) evidence.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

The central claims rest on assumptions about LLM output quality, self-report validity, and ranking aggregation; no independent benchmark or control condition is provided. Free parameters are limited but include hand-tuned prompts, an unspecified ranking-score confidence modification, and manual coding weights. No new physical or theoretical entities are introduced; 'Dynamic Surveys' is an engineering artifact, not a postulated entity.

free parameters (3)
  • LLM clustering and follow-up prompts = Appendix A.1/A.2
    Hand-tuned in a protostudy (Section 3.2); central to the system output, with no version control or systematic ablation.
  • Borda score confidence-bound parameter = not specified
    Section 3.6: ranking scores use a 'slight modification' of the Borda rule with a lower confidence bound; the exact formula or confidence level is not given.
  • Cluster accuracy coding weights = 1, 0.5, 0
    Section 5.1.2: manual rubric for coding response fit; weights were chosen by the researchers and no inter-rater reliability is reported.
axioms (4)
  • domain assumption LLM outputs (clusters, follow-up questions, summaries) are sufficiently accurate and useful for qualitative insight generation.
    The entire platform depends on GPT-4o; accuracy is only manually scored on two datasets (average 88.9) with no inter-rater reliability (Section 5.1.2).
  • domain assumption Respondents' self-reports on the post-study survey accurately reflect engagement and data-quality effects.
    No behavioral metrics or control conditions; 44 of 93 self-selected respondents; acquiescence bias acknowledged in Section 7.
  • domain assumption The Borda rule with a lower-confidence-bound modification is an appropriate aggregation of cluster rankings.
    Adopted without comparison to alternatives; exact formula unspecified (Section 3.6).
  • domain assumption Participants can meaningfully compare Dynamic Surveys to 'other survey tools' from memory.
    The post-study survey asks about comparison without a baseline or control condition (Section 4.2).

reviewed 2026-08-04 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Dynamic Surveys: Using LLMs to Blend Qualitative Depth,Quantitative Structure, and Collaborative Interaction." pith.science (2026). https://pith.science/paper/TMKJ5L3J

@misc{pith2026260800357,
  author       = {Pith},
  title        = {Pith review of: Dynamic Surveys: Using LLMs to Blend Qualitative Depth,Quantitative Structure, and Collaborative Interaction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TMKJ5L3J}},
  note         = {Machine review of arXiv:2608.00357}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Surveys are a powerful tool for collecting data and eliciting insights on social phenomena, and are critical in product design, marketing, scientific research. However, traditional open-ended and closed-ended question formats limit researchers' ability to capture data that combines both the richness of qualitative insights and the analytical rigor of quantitative data. To address these problems, we propose Dynamic Surveys, a survey platform that uses Large Language Models (LLMs) to dynamically cluster qualitative responses in real time and to elicit quantitative ratings and rankings on those clusters and qualitative reflections on how their views compare to broader respondent trends, especially helpful in early-stage or exploratory research settings. This process generates a report showing survey creators and respondents the clustered responses as well as each cluster's rank, rating distribution, and follow-up reflections. To evaluate Dynamic Surveys, we conducted two field studies with 93 participants over a 2-month period. In the first study, 52 students provided input for a career workshop, while in the second, 41 students gave feedback on gaps in their academic curriculum. Of these, 44 respondents filled out a survey on their experience using Dynamic Surveys. We also shared the generated report with 4 individuals who were interested in the insights for their work, and interviewed them to understand their perspectives on the results and any contextual risks they saw in the platform design. Our findings suggest that Dynamic Surveys not only provide richer and deeper insights into responses compared with traditional survey tools, but also increase engagement and foster a sense of community. We discuss broader implications for the design of survey platforms that blend qualitative depth with quantitative structure, facilitating richer insights and offering more collaborative interactions.

Figures

Figures reproduced from arXiv: 2608.00357 by Aidan Ladenburg, Ansh Kumar, David T. Lee, Dishita Jhawar, Ipsita Bisht, Kehua Lei, Zahra Petiwala, Zili Wang.

Figure 1
Figure 1. Figure 1: Workflow of Dynamic Surveys. The green flow at the bottom illustrates the respondent flow during the [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Screenshot of the first screen of the Career Workshop Survey, displaying cluster rankings, each cluster’s [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Screenshot from the Department Course Curriculum Survey showing details of the first cluster. It [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: (a) Respondents’ answers to the post-study survey question, “I felt that the clusters were clear, coherent, [PITH_FULL_IMAGE:figures/full_fig_p013_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: (a) Respondents’ answers to the post-study survey question, “I appreciated being asked to share more [PITH_FULL_IMAGE:figures/full_fig_p017_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: (a) Respondents’ answers to the post-study survey question, “I resonated with many of the responses [PITH_FULL_IMAGE:figures/full_fig_p019_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

86 extracted references · 5 linked inside Pith

  1. [1]

    Matthew Andreotta, Robertus Nugroho, Mark J Hurlstone, Fabio Boschetti, Simon Farrell, Iain Walker, and Cecile Paris. 2019. Analyzing social media data: A mixed-methods framework combining computational and qualitative text analysis.Behavior research methods51 (2019), 1766–1781

  2. [2]

    Aneesha Bakharia, Peter Bruza, Jim Watters, Bhuva Narayan, and Laurianne Sitbon. 2016. Interactive topic modeling for aiding qualitative content analysis. InProceedings of the 2016 ACM on Conference on Human Information Interaction and Retrieval. 213–222

  3. [3]

    Fran Baum, Colin MacDougall, and Danielle Smith. 2006. Participatory action research.Journal of epidemiology and community health60, 10 (2006), 854

  4. [4]

    Eric PS Baumer, Xiaotong Xu, Christine Chu, Shion Guha, and Geri K Gay. 2017. When subjects interpret the data: Social media non-use as a case for adapting the delphi method to cscw. InProceedings of the 2017 acm conference on computer supported cooperative work and social computing. 1527–1543

  5. [5]

    Pat Bazeley. 2009. Analysing qualitative data: More than ‘identifying themes’.Malaysian journal of qualitative research 2, 2 (2009), 6–22

  6. [6]

    2019.Advances in questionnaire design, development, evaluation and testing

    Paul C Beatty, Debbie Collins, Lyn Kaye, Jose-Luis Padilla, Gordon B Willis, and Amanda Wilmot. 2019.Advances in questionnaire design, development, evaluation and testing. John Wiley & Sons

  7. [7]

    2014.The principles of LiquidFeedback

    Jan Behrens, Axel Kistner, Andreas Nitsche, and Björn Swierczek. 2014.The principles of LiquidFeedback. Interacktive Demokratie

  8. [8]

    Quiera S Booker, Jessica D Austin, and Bijal A Balasubramanian. 2021. Survey strategies to increase participant response rates in primary care research studies.Family Practice38, 5 (2021), 699–702

  9. [9]

    William Boone and John Rogan. 2005. Rigour in quantitative analysis: The promise of Rasch analysis techniques. African Journal of Research in Mathematics, Science and Technology Education9, 1 (2005), 25–38

  10. [10]

    Virginia Braun and Victoria Clarke. 2019. Reflecting on reflexive thematic analysis.Qualitative research in sport, exercise and health11, 4 (2019), 589–597

  11. [11]

    Elizabeth A Buchanan and Erin E Hvizdak. 2009. Online survey tools: Ethical and methodological concerns of human research ethics committees.Journal of empirical research on human research ethics4, 2 (2009), 37–48

  12. [12]

    Ashley Castleberry and Amanda Nolen. 2018. Thematic analysis of qualitative research data: Is it as easy as it sounds? Currents in pharmacy teaching and learning10, 6 (2018), 807–815

  13. [13]

    John R Chamberlin and Paul N Courant. 1983. Representative deliberations and representative decisions: Proportional representation and the Borda rule.American Political Science Review77, 3 (1983), 718–733

  14. [14]

    Kathy Charmaz. 2015. Grounded theory.Qualitative psychology: A practical guide to research methods3 (2015), 53–84

  15. [15]

    Nan-Chen Chen, Margaret Drouhard, Rafal Kocielnik, Jina Suh, and Cecilia R Aragon. 2018. Using machine learning to support qualitative coding in social science: Shifting the focus to ambiguity.ACM Transactions on Interactive Intelligent Systems (TiiS)8, 2 (2018), 1–20

  16. [16]

    Nan-chen Chen, Rafal Kocielnik, Margaret Drouhard, Vanessa Peña-Araya, Jina Suh, Keting Cen, Xiangyi Zheng, Cecilia R Aragon, and V Peña-Araya. 2016. Challenges of applying machine learning to qualitative coding. InACM SIGCHI Workshop on Human-Centered Machine Learning

  17. [17]

    Kevin Crowston, Xiaozhong Liu, and Eileen E Allen. 2010. Machine learning and rule-based automated coding of qualitative data.proceedings of the American Society for Information Science and Technology47, 1 (2010), 1–2

  18. [18]

    Brigitte S Cypress. 2019. Data analysis software in qualitative research: Preconceptions, expectations, and adoption. Dimensions of critical care nursing38, 4 (2019), 213–220. 405:24 Kehua Lei et al

  19. [19]

    Shih-Chieh Dai, Aiping Xiong, and Lun-Wei Ku. 2023. LLM-in-the-loop: Leveraging large language model for thematic analysis.arXiv preprint arXiv:2310.15100(2023)

  20. [20]

    Norman K Denzin and Yvonna S Lincoln. 1996. Handbook of qualitative research.Journal of Leisure Research28, 2 (1996), 132

  21. [21]

    Nico Ebert, Björn Scheppler, Kurt Alexander Ackermann, and Tim Geppert. 2023. QButterfly: Lightweight Survey Extension for Online User Interaction Studies for Non-Tech-Savvy Researchers. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–8

  22. [22]

    Sarah Elwood and Helga Leitner. 1998. GIS and community-based planning: Exploring the diversity of neighborhood perspectives and needs.Cartography and Geographic Information Systems25, 2 (1998), 77–88

  23. [23]

    Beatrice Ferrario and Stefanie Stantcheva. 2022. Eliciting people’s first-order concerns: Text analysis of open-ended survey questions. InAEA Papers and Proceedings, Vol. 112. American Economic Association 2014 Broadway, Suite 305, Nashville, TN 37203, 163–169

  24. [24]

    Jessica L Feuston and Jed R Brubaker. 2021. Putting tools in their place: The role of time and perspective in human-AI collaboration for qualitative analysis.Proceedings of the ACM on Human-Computer Interaction5, CSCW2 (2021), 1–25

  25. [25]

    Fábio Freitas, Jaime Ribeiro, Catarina Brandão, Francislê Neri de Souza, António Pedro Costa, and Luís Paulo Reis. 2018. In case of doubt see the manual: A comparative analysis of (self) learning packages qualitative research software. In Computer supported qualitative research: Second international symposium on qualitative research (ISQR 2017). Springer, 176–192

  26. [26]

    Jie Gao, Kenny Tsu Wei Choo, Junming Cao, Roy Ka-Wei Lee, and Simon Perrault. 2023. CoAIcoder: Examining the effectiveness of AI-assisted human-to-human collaboration in qualitative analysis.ACM Transactions on Computer- Human Interaction31, 1 (2023), 1–38

  27. [27]

    Jie Gao, Yuchen Guo, Gionnieve Lim, Tianqin Zhang, Zheng Zhang, Toby Jia-Jun Li, and Simon Tangi Perrault. 2024. CollabCoder: a lower-barrier, rigorous workflow for inductive collaborative qualitative analysis with large language models. InProceedings of the CHI Conference on Human Factors in Computing Systems. 1–29

  28. [28]

    Simret Araya Gebreegziabher, Zheng Zhang, Xiaohang Tang, Yihao Meng, Elena L Glassman, and Toby Jia-Jun Li

  29. [29]

    Matthew Gentzkow, Bryan Kelly, and Matt Taddy. 2019. Text as data.Journal of Economic Literature57, 3 (2019), 535–574

  30. [30]

    Ariel Goldman, Cindy Espinosa, Shivani Patel, Francesca Cavuoti, Jade Chen, Alexandra Cheng, Sabrina Meng, Aditi Patil, Lydia B Chilton, and Sarah Morrison-Smith. 2022. Quad: Deep-learning assisted qualitative data analysis with affinity diagrams. InCHI Conference on Human Factors in Computing Systems Extended Abstracts. 1–7

  31. [31]

    Ashley K Griggs, Marcus E Berzofsky, Bonnie E Shook-Sa, Christine H Lindquist, Kimberly P Enders, Christopher P Krebs, Michael Planty, and Lynn Langton. 2018. The impact of greeting personalization on prevalence estimates in a survey of sexual assault victimization.Public Opinion Quarterly82, 2 (2018), 366–378

  32. [32]

    Justin Grimmer and Brandon M Stewart. 2013. Text as data: The promise and pitfalls of automatic content analysis methods for political texts.Political analysis21, 3 (2013), 267–297

  33. [33]

    Timothy C Guetterman, Tammy Chang, Melissa DeJonckheere, Tanmay Basu, Elizabeth Scruggs, and VG Vinod Vydiswaran. 2018. Augmenting qualitative text analysis with natural language processing: methodological study. Journal of medical Internet research20, 6 (2018), e231

  34. [34]

    Oluwatoyin Hannah Ajilore, Lauretta Eloho Malaka, Aderonke Busayo Sakpere, and Ayomiposi Grace Oluwadebi

  35. [35]

    Harris Héritier, Chloé Allémann, Oleksandr Balakiriev, Victor Boulanger, Sean F Carroll, Noé Froidevaux, Germain Hugon, Yannis Jaquet, Djilani Kebaili, Sandra Riccardi, et al . 2023. Food & You: A digital cohort on personalized nutrition.PLOS Digital Health2, 11 (2023), e0000389

  36. [36]

    Matt-Heun Hong, Lauren A Marsh, Jessica L Feuston, Janet Ruppert, Jed R Brubaker, and Danielle Albers Szafir. 2022. Scholastic: Graphical human-AI collaboration for inductive and interpretive text analysis. InProceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. 1–12

  37. [37]

    Kévin Huguenin, Igor Bilogrevic, Joana Soares Machado, Stefan Mihaila, Reza Shokri, Italo Dacosta, and Jean-Pierre Hubaux. 2017. A predictive model for user motivation and utility implications of privacy-protection mechanisms in location check-ins.IEEE Transactions on Mobile Computing17, 4 (2017), 760–774

  38. [38]

    Jialun Aaron Jiang, Kandrea Wade, Casey Fiesler, and Jed R Brubaker. 2021. Supporting serendipity: Opportunities and challenges for Human-AI Collaboration in qualitative analysis.Proceedings of the ACM on Human-Computer Interaction5, CSCW1 (2021), 1–23

  39. [39]

    Ankur Joshi, Saket Kale, Satish Chandel, and D Kumar Pal. 2015. Likert scale: Explored and explained.British journal of applied science & technology7, 4 (2015), 396–403. Dynamic Surveys 405:25

  40. [40]

    Awais Hameed Khan, Hiruni Kegalle, Rhea D’Silva, Ned Watt, Daniel Whelan-Shamy, Lida Ghahremanlou, and Liam Magee. 2024. Automating Thematic Analysis: How LLMs Analyse Controversial Topics.arXiv preprint arXiv:2405.06919 (2024)

  41. [41]

    Elizabeth Anne Kinsella et al. 2006. Hermeneutics and critical hermeneutics: Exploring possibilities within the art of interpretation. InForum Qualitative Sozialforschung/Forum: Qualitative Social Research, Vol. 7

  42. [42]

    Jan-Christoph Klie, Michael Bugert, Beto Boullosa, Richard Eckart De Castilho, and Iryna Gurevych. 2018. The inception platform: Machine-assisted and knowledge-oriented interactive annotation. InProceedings of the 27th international conference on computational linguistics: System demonstrations. 5–9

  43. [43]

    Travis Kriplean, Jonathan Morgan, Deen Freelon, Alan Borning, and Lance Bennett. 2012. Supporting reflective public thought with considerit. InProceedings of the ACM 2012 conference on Computer Supported Cooperative Work. 265–274

  44. [44]

    Christophe Lejeune. 2011. From normal business to financial crisis... and back again. An illustration of the benefits of Cassandre for qualitative analysis. InForum: Qualitative Sozialforschung, Vol. 12. Institut fur Klinische Sychologie and Gemeindesychologie, Germany

  45. [45]

    Robert P Lennon, Robbie Fraleigh, Lauren J Van Scoy, Aparna Keshaviah, Xindi C Hu, Bethany L Snyder, Erin L Miller, William A Calo, Aleksandra E Zgierska, and Christopher Griffin. 2021. Developing and testing an automated qualitative assistant (AQUA) to support qualitative analysis.Family medicine and community health9, Suppl 1 (2021)

  46. [46]

    Zhuofan Li, Daniel Dohan, and Corey M Abramson. 2021. Qualitative coding in the computational era: A hybrid ap- proach to improve reliability and reduce effort for coding ethnographic interviews.Socius7 (2021), 23780231211062345

  47. [47]

    Megh Marathe and Kentaro Toyama. 2018. Semi-automated coding for qualitative research: A user-centered inquiry and initial prototypes. InProceedings of the 2018 CHI conference on human factors in computing systems. 1–12

  48. [48]

    2013.Qualitative research design: An interactive approach: An interactive approach

    Joseph A Maxwell. 2013.Qualitative research design: An interactive approach: An interactive approach. sage

  49. [49]

    Bryce McDavitt, Laura M Bogart, Matt G Mutchler, Glenn J Wagner, Harold D Green Jr, Sean Jamar Lawrence, Kieta D Mutepfa, and Kelsey A Nogg. 2016. Dissemination as dialogue: building trust and sharing research findings through community engagement.Preventing Chronic Disease13 (2016), E38

  50. [50]

    Raj Mehta and Eugene Sivadas. 1995. Comparing response rates and response content in mail versus electronic mail surveys.Market Research Society. Journal.37, 4 (1995), 1–12

  51. [51]

    Andras Molnar. 2019. SMARTRIQS: A simple method allowing real-time respondent interaction in Qualtrics surveys. Journal of Behavioral and Experimental Finance22 (2019), 161–169

  52. [52]

    Michael J Montoya and Erin E Kent. 2011. Dialogical action: Moving from community-based to community-driven participatory research.Qualitative Health Research21, 7 (2011), 1000–1011

  53. [53]

    Francisco Muñoz-Leiva, Juan Sánchez-Fernández, Francisco Montoro-Ríos, and José Ángel Ibáñez-Zapata. 2010. Improving the response rate and quality in Web-based surveys through the personalization and frequency of reminder mailings.Quality & Quantity44 (2010), 1037–1052

  54. [54]

    Norio Okada, Liping Fang, and D Marc Kilgour. 2013. Community-based decision making in Japan.Group Decision and Negotiation22 (2013), 45–52

  55. [55]

    Lorraine Parker. 1992. Collecting data the e-mail way.Training & Development46, 7 (1992), 52–55

  56. [56]

    Pol.is. 2015. Pol.is. https://pol.is Collaborative opinion-gathering platform

  57. [57]

    Ulf-Dietrich Reips and Frederik Funke. 2008. Interval-level measurement with visual analogue scales in Internet-based research: VAS Generator.Behavior research methods40, 3 (2008), 699–704

  58. [58]

    Jungwook Rhim, Minji Kwak, Yeaeun Gong, and Gahgene Gweon. 2022. Application of humanization to survey chatbots: Change in chatbot perception, interaction experience, and survey data quality.Computers in Human Behavior 126 (2022), 107034

  59. [59]

    Tim Rietz and Alexander Maedche. 2021. Cody: An AI-based system to semi-automate coding for qualitative research. InProceedings of the 2021 CHI conference on human factors in computing systems. 1–14

  60. [60]

    Jessie Rouder, Olivia Saucier, Rachel Kinder, and Matt Jans. 2021. What to do with all those open-ended responses? Data visualization techniques for survey researchers.Survey Practice(2021)

  61. [61]

    Matthew J Salganik and Karen EC Levy. 2015. Wiki surveys: Open and quantifiable social data collection.PloS one10, 5 (2015), e0123483

  62. [62]

    Daniel Schugurensky and Laurie Mook. 2024. Participatory budgeting and local development: Impacts, challenges, and prospects.Local Development & Society5, 3 (2024), 433–445

  63. [63]

    Pontus Stenetorp, Sampo Pyysalo, Goran Topić, Tomoko Ohta, Sophia Ananiadou, and Jun’ichi Tsujii. 2012. BRAT: a web-based tool for NLP-assisted text annotation. InProceedings of the Demonstrations at the 13th Conference of the European Chapter of the Association for Computational Linguistics. 102–107

  64. [64]

    Samantha L Thomas, Hannah Pitt, Simone McCarthy, Grace Arnot, and Marita Hennessy. 2024. Methodological and practical guidance for designing and conducting online qualitative surveys in public health.Health Promotion International39, 3 (2024), daae061. 405:26 Kehua Lei et al

  65. [65]

    Lev Velykoivanenko, Kavous Salehzadeh Niksirat, Stefan Teofanovic, Bertil Chapuis, Michelle L Mazurek, and Kévin Huguenin. 2024. Designing a Data-Driven Survey System: Leveraging Participants’ Online Data to Personalize Surveys. InProceedings of the CHI Conference on Human Factors in Computing Systems. 1–22

  66. [66]

    James Wen and Ashley Colley. 2022. Hybrid Online Survey System with Real-Time Moderator Chat. InProceedings of the 21st International Conference on Mobile and Ubiquitous Multimedia. 257–258

  67. [67]

    Ryan Wesslen. 2018. Computer-assisted text analysis for social science: Topic models and beyond.arXiv preprint arXiv:1803.11045(2018)

  68. [68]

    Justin Wolfers and Eric Zitzewitz. 2004. Prediction markets.Journal of economic perspectives18, 2 (2004), 107–126

  69. [69]

    Kevin B Wright. 2005. Researching Internet-based populations: Advantages and disadvantages of online survey research, online questionnaire authoring software packages, and web survey services.Journal of computer-mediated communication10, 3 (2005), JCMC1034

  70. [70]

    Ziang Xiao, Michelle X Zhou, Q Vera Liao, Gloria Mark, Changyan Chi, Wenxi Chen, and Huahai Yang. 2020. Tell me about yourself: Using an AI-powered chatbot to conduct conversational surveys with open-ended questions.ACM Transactions on Computer-Human Interaction (TOCHI)27, 3 (2020), 1–37

  71. [71]

    Jasy Liew Suet Yan, Nancy McCracken, and Kevin Crowston. 2014. Semi-automatic content analysis of qualitative data.IConference 2014 Proceedings(2014)

  72. [72]

    Erica C Yu, Scott Fricker, and Brandon Kopp. 2015. Can survey instructions relieve respondent burden. In70th Annual Conference of the American Association for Public Opinion Research, Hollywood, FL

  73. [73]

    Brahim Zarouali, Theo Araujo, Jakob Ohme, and Claes de Vreese. 2024. Comparing chatbots and online surveys for (longitudinal) data collection: an investigation of response characteristics, data quality, and user evaluation. Communication Methods and Measures18, 1 (2024), 72–91

  74. [74]

    He Zhang, Chuhao Wu, Jingyi Xie, Yao Lyu, Jie Cai, and John M Carroll. 2023. Redefining qualitative analysis in the AI era: Utilizing ChatGPT for efficient thematic analysis.arXiv preprint arXiv:2309.10771(2023)

  75. [75]

    He Zhang, Chuhao Wu, Jingyi Xie, Fiona Rubino, Sydney Graver, ChanMin Kim, John M Carroll, and Jie Cai. 2024. When Qualitative Research Meets Large Language Model: Exploring the Potential of QualiGPT as a Tool for Qualitative Coding.arXiv preprint arXiv:2407.14925(2024). A Prompts A.1 Clustering prompt I’m running a survey that involves qualitative data a...

  76. [78]

    You must strike a balance and create a theme that is clear, concise, and useful for thematic analysis

    Themes are general categories that represent shared ideas between responses. You must strike a balance and create a theme that is clear, concise, and useful for thematic analysis

  77. [79]

    Someone should be able to understand the core aspect of the theme by only reading the theme name

    Theme names should be at least 3 words long. Someone should be able to understand the core aspect of the theme by only reading the theme name

  78. [80]

    Redundant themes are not useful and responses should be placed in existing themes before new themes are generated

    Any themes generated should be fully unique with no overlap between them. Redundant themes are not useful and responses should be placed in existing themes before new themes are generated

  79. [81]

    Responses can be in 1 or many themes

  80. [82]

    I’d like you to think step-by-step using the following strategy: Dynamic Surveys 405:27

    Try to minimize the number of themes generated while still capturing the IMPORTA- NT aspects of the response. I’d like you to think step-by-step using the following strategy: Dynamic Surveys 405:27

Showing first 80 references.

This paper was first reviewed by deepseek-v4-flash on August 4, 2026.