Pith. sign in

REVIEW 4 major objections 4 minor 52 references

Mic Drop or Data Flop? Evaluating the Fitness for Purpose of AI Voice Interviewers for Data Collection within Quantitative & Qualitative Research Contexts

T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read A review of field evidence concludes AI voice interviewers already beat IVR for surveys, but only partly fit qualitative work.

desk verdict A useful framework, but the central IVR comparison is inferred, not measured, and one core study is the authors' own. read the letter →

arxiv 2509.01814 v1 pith:IEUAIZEO submitted 2025-09-01 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords AIinterviewersvoice-basedsurveysinteractivevoiceresponseautomaticspeechrecognitionlargelanguagemodelsqualitativedatacollectionquantitativesurveymethodology
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks when an AI interviewer—a voice agent built from a speech recognizer, a large language model, and a text-to-speech engine—can replace or supplement traditional phone surveys. Reviewing early field evidence, it argues that these systems already outperform interactive voice response (IVR) on the two things a survey needs: faithfully receiving and recording answers, and making sensible mid-interview decisions such as clarifying, repeating, or branching. For structured quantitative questionnaires, the paper concludes, AI interviewers are fit for purpose today. For qualitative interviews built on open-ended answers, follow-up probes, and emotional cues, it finds the technology only sometimes works: real-time transcription errors, weak emotion detection, and uneven probe quality make adoption context-dependent. The paper's practical point is that researchers should pick AI interviewers where probing depth is not pivotal and post-processing is possible, and keep human interviewers where rapport and emotional nuance carry the data.

What carries the argument

The load-bearing object is the AI interviewer architecture that chains three components: automatic speech recognition (ASR) to transcribe what respondents say, a transformer-based large language model to interpret the transcript and choose the next utterance, and text-to-speech (TTS) to voice that utterance. This pipeline is what converts a fixed decision tree into a conversational agent that can branch, clarify, repeat, and probe. The paper evaluates both IVR and this pipeline along two dimensions borrowed from survey-quality work: input/output performance (can the system faithfully receive and deliver answers?) and verbal reasoning (can it vary the conversation intelligently?). The same ar

What would settle it

A preregistered randomized field experiment on a general-population panel assigns the same survey to an AI interviewer, an IVR line, and a trained human interviewer; if the AI interviewer does not beat IVR on completion, data-quality metrics, and branching accuracy, or if blind coders rate its open-ended probes as no better than scripted follow-ups, the paper's central ordering fails.

Watch

Extended reading notes

Core claim

The core claim is a capability ordering with a boundary. Drawing on three independent field studies completed in the last year, the paper argues that LLM-based AI interviewers already clear the bar for quantitative data collection: they can administer long instruments with conditional branching, randomize question and answer order, re-prompt on out-of-range or ambiguous answers, and ask questions in their original wording, none of which IVR does well. For qualitative data collection, the same evidence supports a qualified advance rather than a clean win: AI interviewers can ask open-ended questions and generate unscripted follow-ups, but streaming speech-to-text error rates near 10 percent a

Load-bearing premise

The conclusion assumes that the few early field studies of AI interviewers, based mainly on student samples and too small for subgroup analysis, are representative enough to establish where the technology already outperforms IVR.

Editorial extensions

If this is right

  • Survey researchers can deploy AI interviewers for structured quantitative questionnaires now, including instruments with branching logic and clarifications, without rewriting question wording to fit the medium.
  • For open-ended qualitative work, the default should be to include AI interviewers only when post-hoc transcription of recordings is available to repair streaming errors and when the study design does not hinge on deep probing.
  • Sensitive-topic surveys are a likely early niche: automation reduces social desirability bias, and the paper's field evidence suggests respondents accept AI interviewers, even preferring them for some sensitive questions.
  • As long as emotion detection and expression remain unreliable, AI interviewers will not replace human interviewers for rapport-driven qualitative methods such as in-depth interviews, focus groups, or cognitive pretesting.
  • The evidence base is still too narrow to support claims about specific subpopulations, because most respondents studied so far are students and sample sizes are small.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the text-passing architecture itself is the limiting factor, the arrival of low-latency speech-to-speech models that skip the intermediate transcript could shrink the paper's main caveats—emotion and transcription—faster than its review horizon suggests.
  • The paper's comparison is implicitly against IVR as a baseline; against human interviewers the evidence is thinner, so the actionable reading may be 'better than the automated predecessor' rather than 'ready to replace people'.
  • A testable extension would log streaming ASR errors per open-ended response and correlate error counts with response length; the paper's reasoning predicts error rates rise as answers lengthen, which would formalize why qualitative surveys are more vulnerable.
  • The social desirability benefit is imported largely from earlier automated-interview literature; a randomized mode experiment comparing AI interviewer, IVR, web, and human interviews on the same sensitive items would determine whether that advantage is a modality effect or a general automation effect.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. Tirumala et al. present a position paper evaluating AI voice interviewers (ASR + LLM + TTS) against IVR for survey data collection. They propose two evaluation dimensions—input/output performance and verbal reasoning—and apply them separately to IVR (Section 3.1) and AI interviewers (Section 3.2), drawing on three recent field studies by Leybzon et al. (2025), Lang & Eskenazi (2025), and Wuttke et al. (2024) plus broader ML literature. The paper concludes that AI interviewers already exceed IVR capabilities for both quantitative and qualitative data collection, while noting that real-time ASR error (~10.9%), limited emotion detection, and uneven follow-up quality make adoption context-dependent. It closes with a future research agenda on respondent representativeness, preferences, and methodological variation.

Significance. The paper offers a useful conceptual framework for comparing automated interview systems, and it is honest about several open problems (sample representativeness, follow-up quality, latency). It does not present new experimental data, machine-checked proofs, or code; its strength is as a synthesis. If the conclusions were better supported, this would be an agenda-setting contribution for survey methodology. Currently, the evidence base is not strong enough for the 'already exceed' claim, and the paper's own caveats and conflict considerations need to be addressed before publication.

major comments (4)
  1. [Abstract; §3.1–§3.2; §4] The central claim that AI interviewers 'already exceed IVR capabilities' is a cross-study inference, not a measured result. Section 3.1 evaluates IVR from older literature (Kim et al., 2011; Dillman et al., 2009; Amaya, 2021) and Section 3.2 evaluates AI interviewers from three separate field studies, none of which includes an IVR control arm. Because the technologies differ in response modality, latency, and error patterns, the net data-quality difference is unidentified. The paper itself reports ~10.9% real-time ASR word error (Section 3.2.1), which may be a source of measurement error that touch-tone IVR avoids for closed-ended questions. Please either temper the conclusion to 'may offer additional functionality', add an explicit comparison that makes the indirect nature of the inference clear, or cite a direct head-to-head study.
  2. [§3.2, references (Leybzon et al., 2025)] One of the three core field studies is authored by the paper's own team: Leybzon et al. (2025) shares authors Tirumala, Jain, and Leybzon with this manuscript. This study is used as central evidence for AI interviewer capabilities (e.g., the 123-question SSRS panel and interruption handling in §3.2.2/§3.2.1). The manuscript nowhere discloses this conflict of interest and even describes the three studies as 'independently' conducted. At minimum, the author relationship must be disclosed and the argument should be assessed with this study removed; the qualitative conclusions should not rely on an unacknowledged self-supplied data point.
  3. [§3.2.2; §4] The qualitative data collection claim is extrapolated from evidence that is not about AI voice interviewers. The three field studies are primarily quantitative survey administrations; the follow-up quality evidence cited in §3.2.2 comes from text-based LLM chatbots (Cuevas et al., 2024; Kuric et al., 2024; Geisen, 2024; Chopra & Haaland, 2023) and from one semi-structured interview generation study (Zeng et al., 2023) that is not a field deployment. The paper honestly notes the conflicting results, but the abstract's binary claim that AI interviewers 'already exceed IVR... for qualitative data collection' is stronger than this indirect evidence supports. Please label this as an extrapolation and, if possible, restrict the qualitative claim to 'may be capable of richer probing than IVR, pending direct tests'.
  4. [§4.1] The generalizability caveats in §4.1 undercut the abstract's blanket wording. The paper concedes that most studies use student samples and that Leybzon et al.'s probability-based panel is underpowered for subpopulation inference. The conclusion 'AI interviewers already exceed IVR' is stated without these limits. A revised version should carry the population limitation into the abstract and conclusion (e.g., 'in the convenience samples studied to date').
minor comments (4)
  1. [§3.2.1] Typographical: '˜10.9%' should be '~10.9%'.
  2. [References] Several author names have garbled diacritics (e.g., 'D´efossez', 'Mazar´e', 'Am´elie'); please use proper encoding.
  3. [§1.2.2] Figure 1 is referenced but no actual figure appears in the manuscript text; please ensure the architecture diagram is included.
  4. [§3.1.1] The example quote from Amaya (2021) has slightly awkward wording ('For a woman, press two'); verify exact wording as it appears in the source.

Circularity Check

0 steps flagged · score 2.0 of 10

No definitional circularity; the central comparison is an inference across separate literatures and one core field study is self-authored, but the conclusion is not forced by construction or self-citation alone.

full rationale

The paper's claimed derivation chain is a capability review, not a formal derivation. Its definitions of IVR (§1.2.1) and AI interviewer (§1.2.2) are structural and do not smuggle in the conclusion that AI interviewers exceed IVR; the verbal-reasoning advantage follows from the definitions, but the paper treats that as a framing, not an empirical prediction. The central claim that 'AI interviewers already exceed IVR capabilities' is based on three field studies (§3.2). One of these, Leybzon et al. (2025), is by three of the present authors and is described as 'independent' and as part of the 'core of our analysis' (§3.2), which is an undeclared conflict of interest. However, Lang & Eskenazi (2025) and Wuttke et al. (2024) are independent and support the same capability conclusions, so the claim does not reduce to the self-citation. The 'exceed IVR' conclusion is an inference across separate literatures because no cited study includes an IVR control arm; this is an evidentiary gap, not a circular step. The paper itself flags limits (respondent representativeness, sample size) in §4.1. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from prior work, and no known result is merely relabeled. Score 2 reflects one non-load-bearing self-citation that should have been disclosed; the derivation itself is not circular.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters or invented entities. The analysis rests on domain assumptions about the architecture of AI interviewers, the validity and independence of the cited field studies, the role of emotion in qualitative interviewing, and the comparative relevance of IVR limitations. These assumptions are reasonable but not empirically established here.

assumptions (4)
  • domain assumption The ASR + LLM + TTS pipeline is the standard architecture for contemporary AI interviewers, and speech-to-speech models are not yet feasible in real-time.
    Section 1.2.2 defines AI interviewers through this architecture and excludes speech-to-speech models due to latency; the capability assessment is scoped by this assumption.
  • domain assumption The cited field studies (Leybzon et al. 2025; Lang & Eskenazi 2025; Wuttke et al. 2024) are valid, representative, and independent evidence for AI interviewer performance.
    Section 3.2 states these works 'form the core of our analysis'; no systematic inclusion criteria are given, and one study is by the authors.
  • domain assumption Emotion detection and expression are required for high-quality qualitative data collection.
    Sections 3.2.1 and 3.2.2 argue that limited emotion handling reduces qualitative fitness; this is a substantive assumption about qualitative interviewing, not derived from data in the paper.
  • domain assumption IVR systems cannot conduct dynamic follow-ups, so any system that can is preferable for qualitative collection.
    Section 3.1.2 states IVR lacks dynamic response capabilities; the conclusion that AI interviewers 'exceed' IVR for qualitative data follows partly from this framing rather than from a direct data-quality comparison.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Mic Drop or Data Flop? Evaluating the Fitness for Purpose of AI Voice Interviewers for Data Collection within Quantitative & Qualitative Research Contexts." pith.science (2026). https://pith.science/paper/IEUAIZEO

@misc{pith2026250901814,
  author       = {Pith},
  title        = {Pith review of: Mic Drop or Data Flop? Evaluating the Fitness for Purpose of AI Voice Interviewers for Data Collection within Quantitative & Qualitative Research Contexts},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IEUAIZEO}},
  note         = {Machine review of arXiv:2509.01814}
}
read the original abstract

Transformer-based Large Language Models (LLMs) have paved the way for "AI interviewers" that can administer voice-based surveys with respondents in real-time. This position paper reviews emerging evidence to understand when such AI interviewing systems are fit for purpose for collecting data within quantitative and qualitative research contexts. We evaluate the capabilities of AI interviewers as well as current Interactive Voice Response (IVR) systems across two dimensions: input/output performance (i.e., speech recognition, answer recording, emotion handling) and verbal reasoning (i.e., ability to probe, clarify, and handle branching logic). Field studies suggest that AI interviewers already exceed IVR capabilities for both quantitative and qualitative data collection, but real-time transcription error rates, limited emotion detection abilities, and uneven follow-up quality indicate that the utility, use and adoption of current AI interviewer technology may be context-dependent for qualitative data collection efforts.

Figures

Figures reproduced from arXiv: 2509.01814 by the authors.

Figure 1
Figure 1. Standard AI interviewer architecture As depicted in [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

52 extracted references · 41 canonical work pages

  1. [1]

    Conducting semi-structured interviews

    William C Adams. Conducting semi-structured interviews. Handbook of practical program evaluation, pp.\ 492--505, 2015

  2. [2]

    A comprehensive evaluation of incremental speech recognition and diarization for conversational ai

    Angus Addlesee, Yanchao Yu, and Arash Eshghi. A comprehensive evaluation of incremental speech recognition and diarization for conversational ai. In Proceedings of the 28th International Conference on Computational Linguistics, pp.\ 3492--3503, 2020

  3. [3]

    Speech emotion recognition in conversations using artificial intelligence: a systematic review and meta-analysis

    Ghada Alhussein, Ioannis Ziogas, Shiza Saleem, and Leontios J Hadjileontiadis. Speech emotion recognition in conversations using artificial intelligence: a systematic review and meta-analysis. Artificial Intelligence Review, 58 0 (7): 0 198, 2025

  4. [4]

    How call-in options affect address-based web surveys

    Ashley Amaya. How call-in options affect address-based web surveys. Pew Re-search Center, 2021

  5. [5]

    Ai-assisted conversational interviewing: Effects on data quality and user experience

    Soubhik Barari, Jarret Angbazo, Natalie Wang, Leah M Christian, Elizabeth Dean, Zoe Slowinski, and Brandon Sepulvado. Ai-assisted conversational interviewing: Effects on data quality and user experience. arXiv preprint arXiv:2504.13908, 2025

  6. [6]

    The utility of probability-based online surveys: a literature review

    Oriol J Bosch and Olga Maslovskaya. The utility of probability-based online surveys: a literature review. In 2023 AAPOR Webinar Series. AAPOR, 2023

  7. [7]

    Gpt, pretend you are a survey researcher: Results from a systematic literature review exploring the use of large language models within survey research

    Trent D Buskirk, Adam Eck, Leah von der Heyde, and Florian Keusch. Gpt, pretend you are a survey researcher: Results from a systematic literature review exploring the use of large language models within survey research. In 80th Annual AAPOR Conference. AAPOR, 2025

  8. [8]

    Conducting qualitative interviews with ai

    Felix Chopra and Ingar Haaland. Conducting qualitative interviews with ai. CESifo Working Paper 10666, Center for Economic Studies and ifo Institute (CESifo), Munich, September 2023. URL https://www.ifo.de/DocDL/cesifo1_wp10666.pdf

Show all 52 references
  1. [9]

    Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models, 2023

    Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou. Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models, 2023. URL https://arxiv.org/abs/2311.07919

  2. [10]

    Automating telephone surveys: using t-acasi to obtain data on sensitive topics

    Philip C Cooley, Heather G Miller, James N Gribble, and Charles F Turner. Automating telephone surveys: using t-acasi to obtain data on sensitive topics. Computers in Human Behavior, 16 0 (1): 0 1--11, 2000

  3. [11]

    Interactive voice response: review of studies 1989--2000

    Ross Corkrey and Lynne Parkinson. Interactive voice response: review of studies 1989--2000. Behavior Research Methods, Instruments, & Computers, 34 0 (3): 0 342--353, 2002

  4. [12]

    Scurrell, Eva M

    Alejandro Cuevas, Jennifer V. Scurrell, Eva M. Brown, Jason Entenmann, and Madeleine I. G. Daepp. Collecting qualitative data at scale with large language models: A case study, 2024. URL https://arxiv.org/abs/2309.10187

  5. [13]

    Recent advances in speech language models: A survey, 2025

    Wenqian Cui, Dianzhi Yu, Xiaoqi Jiao, Ziqiao Meng, Guangyan Zhang, Qichao Wang, Yiwen Guo, and Irwin King. Recent advances in speech language models: A survey, 2025. URL https://arxiv.org/abs/2410.03751

  6. [14]

    Simsensei kiosk: A virtual human interviewer for healthcare decision support

    David DeVault, Ron Artstein, Grace Benn, Teresa Dey, Ed Fast, Alesia Gainer, Kallirroi Georgila, Jon Gratch, Arno Hartholt, Margaux Lhommet, et al. Simsensei kiosk: A virtual human interviewer for healthcare decision support. In Proceedings of the 2014 international conference...

  7. [15]

    Response rate and measurement differences in mixed-mode surveys using mail, telephone, interactive voice response (ivr) and the internet

    Don A Dillman, Glenn Phelps, Robert Tortora, Karen Swift, Julie Kohrell, Jodi Berck, and Benjamin L Messer. Response rate and measurement differences in mixed-mode surveys using mail, telephone, interactive voice response (ivr) and the internet. Social science research, 38 0 (...

  8. [16]

    Moshi: a speech-text foundation model for real-time dialogue, 2024

    Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour. Moshi: a speech-text foundation model for real-time dialogue, 2024. URL https://arxiv.org/abs/2410.00037

  9. [17]

    We need to talk: Audio surveys and information extraction

    Vincenzo Galasso, Tommaso Nannicini, and Debora Nozza. We need to talk: Audio surveys and information extraction. Technical report, CESifo Working Paper, 2024

  10. [18]

    Prompting insight: Enhancing open-ended survey responses with ai-powered follow-ups

    Emily Geisen. Prompting insight: Enhancing open-ended survey responses with ai-powered follow-ups. In 79th Annual AAPOR Conference. AAPOR, 2024

  11. [19]

    Digging deep: How diverse identities are associated with different survey mode preferences for adverse, traumatic experiences

    Ris \"e B Goldstein, Megan A Hendrich, Randall K Thomas, and Robert Petrin. Digging deep: How diverse identities are associated with different survey mode preferences for adverse, traumatic experiences. In 80th Annual AAPOR Conference. AAPOR, 2025

  12. [20]

    The audio check: A method for improving data quality and detecting data fabrication

    Robin Gomila, Rebecca Littman, Graeme Blair, and Elizabeth Levy Paluck. The audio check: A method for improving data quality and detecting data fabrication. Social Psychological and Personality Science, 8 0 (4): 0 424--433, 2017

  13. [21]

    Interview mode and measurement of sexual behaviors: Methodological issues

    James N Gribble, Heather G Miller, Susan M Rogers, and Charles F Turner. Interview mode and measurement of sexual behaviors: Methodological issues. Journal of Sex research, 36 0 (1): 0 16--24, 1999

  14. [22]

    Survey methodology

    Robert M Groves, Floyd J Fowler Jr, Mick P Couper, James M Lepkowski, Eleanor Singer, and Roger Tourangeau. Survey methodology. John Wiley & Sons, 2011

  15. [23]

    A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Informatio...

  16. [24]

    Comparative analysis and review of interactive voice response systems

    Itorobong A Inam, Ambrose A Azeta, and Olawande Daramola. Comparative analysis and review of interactive voice response systems. In 2017 Conference on Information Communication Technology and Society (ICTAS), pp.\ 1--6. IEEE, 2017

  17. [25]

    Inherent usability problems in interactive voice response systems

    Hee-Cheol Kim, Deyun Liu, and Ho-Won Kim. Inherent usability problems in interactive voice response systems. In Human-Computer Interaction. Users and Applications: 14th International Conference, HCI International 2011, Orlando, FL, USA, July 9-14, 2011, Proceedings, Part IV 14...

  18. [26]

    Comparing data from chatbot and web surveys: Effects of platform and conversational style on survey response quality

    Soomin Kim, Joonhwan Lee, and Gahgene Gweon. Comparing data from chatbot and web surveys: Effects of platform and conversational style on survey response quality. In Proceedings of the 2019 CHI conference on human factors in computing systems, pp.\ 1--12, 2019

  19. [27]

    Social desirability bias in cati, ivr, and web surveys: The effects of mode and question sensitivity

    Frauke Kreuter, Stanley Presser, and Roger Tourangeau. Social desirability bias in cati, ivr, and web surveys: The effects of mode and question sensitivity. Public opinion quarterly, 72 0 (5): 0 847--865, 2008

  20. [28]

    Measuring the accuracy of automatic speech recognition solutions

    Korbinian Kuhn, Verena Kersken, Benedikt Reuter, Niklas Egger, and Gottfried Zimmermann. Measuring the accuracy of automatic speech recognition solutions. ACM Transactions on Accessible Computing, 16 0 (4): 0 1--23, 2024

  21. [29]

    Unmoderated usability studies evolved: Can gpt ask useful follow-up questions? International Journal of Human--Computer Interaction, pp.\ 1--18, 2024

    Eduard Kuric, Peter Demcak, and Matus Krajcovic. Unmoderated usability studies evolved: Can gpt ask useful follow-up questions? International Journal of Human--Computer Interaction, pp.\ 1--18, 2024

  22. [30]

    Telephone surveys meet conversational ai: Evaluating a llm-based telephone survey system at scale

    Max Lang and Sol Eskenazi. Telephone surveys meet conversational ai: Evaluating a llm-based telephone survey system at scale. In 80th Annual AAPOR Conference. AAPOR, 2025

  23. [31]

    Investigating memory constraints on recall of options in interactive voice response system messages

    Ludovic Le Bigot, Lo \" c Caroux, Christine Ros, Agn \`e s Lacroix, and Val \'e rie Botherel. Investigating memory constraints on recall of options in interactive voice response system messages. Behaviour & Information Technology, 32 0 (2): 0 106--116, 2013

  24. [32]

    The promise & pitfalls of ai‑augmented survey research

    Joshua Lerner. The promise & pitfalls of ai‑augmented survey research. NORC research library, Oct 2024. URL https://www.norc.org/research/library/promise-pitfalls-ai-augmented-survey-research.html. Accessed on June 20, 2024

  25. [33]

    Leybzon, Shreyas Tirumala, Nishant Jain, Summer Gillen, Michael Jackson, Cameron McPhee, and Jennifer Schmidt

    Danny D. Leybzon, Shreyas Tirumala, Nishant Jain, Summer Gillen, Michael Jackson, Cameron McPhee, and Jennifer Schmidt. Ai telephone surveying: Automating quantitative data collection with an ai interviewer. In 80th Annual AAPOR Conference. AAPOR, 2025

  26. [34]

    Social evaluation of text-to-speech voices by adults and children

    Kevin D Lilley, Ellen Dossey, Michelle Cohn, Cynthia G Clopper, Laura Wagner, and Georgia Zellou. Social evaluation of text-to-speech voices by adults and children. Speech Communication, 166: 0 103163, 2025

  27. [35]

    Why do survey respondents disclose more when computers ask the questions? Public opinion quarterly, 77 0 (4): 0 888--935, 2013

    Laura H Lind, Michael F Schober, Frederick G Conrad, and Heidi Reichert. Why do survey respondents disclose more when computers ask the questions? Public opinion quarterly, 77 0 (4): 0 888--935, 2013

  28. [36]

    It’s only a computer: Virtual humans increase willingness to disclose

    Gale M Lucas, Jonathan Gratch, Aisha King, and Louis-Philippe Morency. It’s only a computer: Virtual humans increase willingness to disclose. Computers in Human Behavior, 37: 0 94--100, 2014

  29. [37]

    Application of humanization to survey chatbots: Change in chatbot perception, interaction experience, and survey data quality

    Jungwook Rhim, Minji Kwak, Yeaeun Gong, and Gahgene Gweon. Application of humanization to survey chatbots: Change in chatbot perception, interaction experience, and survey data quality. Computers in Human Behavior, 126: 0 107034, 2022

  30. [38]

    Precision and disclosure in text and voice interviews on smartphones

    Michael F Schober, Frederick G Conrad, Christopher Antoun, Patrick Ehlen, Stefanie Fail, Andrew L Hupp, Michael Johnston, Lucas Vickers, H Yanna Yan, and Chan Zhang. Precision and disclosure in text and voice interviews on smartphones. PloS one, 10 0 (6): 0 e0128337, 2015

  31. [39]

    Emphasis control for parallel neural tts

    Shreyas Seshadri, Tuomo Raitio, Dan Castellani, and Jiangchuan Li. Emphasis control for parallel neural tts. arXiv preprint arXiv:2110.03012, 2021

  32. [40]

    The relationship between interviewer-respondent rapport and data quality

    Hanyu Sun, Frederick G Conrad, and Frauke Kreuter. The relationship between interviewer-respondent rapport and data quality. Journal of Survey Statistics and Methodology, 9 0 (3): 0 429--448, 2021

  33. [41]

    Salmonn: Towards generic hearing abilities for large language models, 2024

    Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. Salmonn: Towards generic hearing abilities for large language models, 2024. URL https://arxiv.org/abs/2310.13289

  34. [42]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurip...

  35. [43]

    Leveraging large language models to power chatbots for collecting user self-reported data

    Jing Wei, Sungdong Kim, Hyunhoon Jung, and Young-Ho Kim. Leveraging large language models to power chatbots for collecting user self-reported data. Proceedings of the ACM on Human-Computer Interaction, 8 0 (CSCW1): 0 1--35, 2024

  36. [44]

    Ai conversational interviewing: Transforming surveys with llms as adaptive interviewers

    Alexander Wuttke, Matthias A enmacher, Christopher Klamm, Max M Lang, Quirin W \"u rschinger, and Frauke Kreuter. Ai conversational interviewing: Transforming surveys with llms as adaptive interviewers. arXiv preprint arXiv:2410.01824, 2024

  37. [45]

    Qwen2.5-omni technical report, 2025

    Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, Bin Zhang, Xiong Wang, Yunfei Chu, and Junyang Lin. Qwen2.5-omni technical report, 2025. URL https://arxiv.org/abs/2503.20215

  38. [46]

    Keeping users engaged during repeated administration of the same questionnaire: Using large language models to reliably diversify questions

    Hye Sun Yun, Mehdi Arjmand, Phillip Sherlock, Michael K Paasche-Orlow, James W Griffith, and Timothy Bickmore. Keeping users engaged during repeated administration of the same questionnaire: Using large language models to reliably diversify questions. arXiv preprint arXiv:2311...

  39. [47]

    Question generation to elicit users’ food preferences by considering the semantic content

    Jie Zeng, Yukiko I Nakano, and Tatsuya Sakato. Question generation to elicit users’ food preferences by considering the semantic content. In Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pp.\ 190--196, 2023

  40. [48]

    Accented text-to-speech synthesis with limited data

    Xuehao Zhou, Mingyang Zhang, Yi Zhou, Zhizheng Wu, and Haizhou Li. Accented text-to-speech synthesis with limited data. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 32: 0 1699--1711, 2024

  41. [49]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

  42. [50]

    @esa (Ref

    \@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...

  43. [51]

    \@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...

  44. [52]

    @open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.