{"id":"dbafabca-e569-4d64-9a84-0ec880d807c3","arxiv_id":"2412.05445","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Adding LLM-based cleanup and an edit-by-prompt agent to a voice-based review app raised users' willingness to share reviews and their self-reported confidence in a 14-participant field study.","lead":"Vocalizer is a phone-based review app that lets people speak their restaurant review and then asks an AI model to clean it up and take instructions for changes. In a 14-person field study, users were more willing to share the AI-enhanced reviews and reported higher confidence in writing reviews, though all measures are self-reported.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Retrospective self-efficacy baseline collected after tool use inflates the reported 4.86-to-9.00 gain; without a pre-intervention baseline the central self-efficacy claim is not established.","rationale":"The reader's weakest assumption identifies the retrospective baseline as a key threat to the self-efficacy claim, and the paper itself acknowledges this limitation in Section 5.3. I agree that this is the single most load-bearing concern: it is not a peripheral design choice but a direct threat to one of the two main quantitative results (Section 4.6) and to the abstract's claim that AI features 'improve users' self-efficacy.' The willingness-to-share result (Section 4.2.1) is less affected by this concern because it is measured per-review throughout the study, not retrospectively. However, the self-efficacy finding is a central contribution, and the retrospective baseline could easily be fixed in a follow-up study by adding the item to the onboarding survey. The proposed test would settle whether the concern lands: if the pre-intervention score is similar to 4.86, the self-efficacy claim survives; if it is higher, the reported effect shrinks. Since the reader already issued a CONDITIONAL verdict with moderate confidence, and this concern supports that conditionality without requiring a stronger rejection, the verdict should remain unchanged. I am not raising objections to the sample size or the lack of released data, as those are already weighed by the reader; the retrospective baseline is the specific decisive flaw that deserves explicit stress-testing.","tokens_in":20935,"tokens_out":3296,"duration_ms":35852,"concrete_test":"Add the same 0-10 self-efficacy item to the onboarding/background questionnaire before any app exposure, administered to all participants. Compare pre-intervention scores with the retrospective 'Unaided' score (4.86) and post-LAV score (9.00) using the same Friedman and Wilcoxon analyses as Section 4.6. If the pre-intervention baseline is not significantly lower than 4.86, the reported self-efficacy gain would be largely an artifact of retrospective bias; if the pre-score is comparable to 4.86, the claim survives this specific threat.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that the LLM-assisted version (LAV) increased users' self-efficacy rests on a comparison between a retrospective 'Unaided' baseline (4.86) and post-intervention scores (LAV 9.00, Overall 9.14). As Section 4.6 and Figure 9 show, the 'Unaided' rating was collected in the final questionnaire, after participants had already used both Vocalizer versions. Section 5.3 explicitly acknowledges this: 'the self-efficacy measurement about the users' prior belief in being able to leave a good review was captured in the final questionnaire, while it may have been more useful to capture it also in the onboarding survey.' This is not a minor measurement detail: the entire self-efficacy contribution—framed in the abstract and conclusion as 'improve users' self-efficacy'—depends on the 4.86 baseline being an accurate pre-intervention estimate. If that retrospective judgment is inflated by the participant's recent successful experience with the tool, the reported jump from 4.86 to 9.00 is partly an artifact, and the causal claim that LAV increased self-efficacy is unsupported. The willingness-to-share result (Section 4.2.1) is a separate, self-report measure and does not suffer from this specific baseline problem, but the self-efficacy finding is a headline contribution and cannot be separated from this flaw. The concern is internal to the paper's design, explicitly acknowledged, and directly threatens the stated conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Vocalizer, a mobile web application for leaving restaurant reviews via voice input, with two versions: a voice-only version (VOV) and an LLM-assisted version (LAV) that adds automatic transcript cleaning, an interactive AI agent, and AI-generated improvement tips. Fourteen participants used both versions over a within-subjects field study lasting up to three weeks, with in-app feedback after each review and post-visit and final questionnaires. The paper reports that users made frequent use of the AI agent (42 of 82 LAV reviews), that willingness to share was significantly higher for LAV than VOV reviews (M = 5.93 vs. 4.53, t(14) = -2.39, p < 0.05), and that self-efficacy rose from an 'Unaided' baseline of 4.86 to 9.00 after LAV use, with a significant Friedman test across four conditions. Qualitative analysis of prompts and open-ended responses is used to characterize user editing strategies and perceptions. The central claims are that LLM-assisted editing increases users' willingness to share and their self-efficacy in writing reviews.","tokens_in":21234,"tokens_out":6193,"duration_ms":64414,"significance":"If the results were robust, this would be a useful contribution to the growing HCI literature on AI-assisted content creation, particularly for mobile and time-constrained review writing. The paper's strengths are its longitudinal within-subjects deployment with a working system, the detailed description of the LLM pipeline and prompts (including full prompt listings in appendices), the analysis of real user-AI interactions, and the candid treatment of several limitations. The qualitative analysis of user prompting strategies is also valuable. However, the quantitative evidence rests on a small sample (N = 14) and on self-report measures whose validity is partly undermined by the study design, especially the retrospectively collected self-efficacy baseline. The self-efficacy finding, which is featured in the abstract and conclusion, is not causally established by the current data.","major_comments":[{"comment":"The self-efficacy baseline is retrospective and is load-bearing for the headline claim. The 'Unaided' score of 4.86 was collected in the final questionnaire, after participants had already used both VOV and LAV; Section 5.3 explicitly states that the 'self-efficacy measurement about the users' prior belief in being able to leave a good review was captured in the final questionnaire.' The key comparison, LAV (9.00) versus Unaided (4.86), therefore contrasts a post-intervention state with a recalled prior state, which is vulnerable to recall bias and effort justification. The paper nevertheless concludes that 'the LAV increased users' self-efficacy in their review writing abilities.' This causal statement is not supported by the data as presented. The authors should either reframe the result as a perceived, retrospectively assessed difference, or provide a genuine pre-intervention baseline (e.g., from the background questionnaire) and re-run the analysis with that baseline.","section":"Section 4.6 and Section 5.3"},{"comment":"The willingness-to-share t-test is reported as 't(14) = -2.39' for what is described as a paired-sample test with 14 participants. For a paired t-test, the degrees of freedom should be 13, not 14. The authors should correct the reported statistic and provide the exact p-value (with df = 13, p is approximately 0.032, so the conclusion remains significant, but the reporting is inaccurate). They should also clarify whether the analysis used per-participant means across the repeated reviews, and consider whether a non-parametric alternative (e.g., Wilcoxon signed-rank) is more appropriate given the small sample and skew evident in the VOV distribution (SD = 1.85). Reporting an effect size and confidence interval would strengthen the paper.","section":"Section 4.2.1"},{"comment":"The study design confounds the automatic LLM-based cleansing of the transcript with the interactive AI agent. The LAV condition includes both the automatic initial improvement (removal of filler words, rephrasing) and the user-facing AI agent, while the VOV condition includes neither. The abstract and conclusion attribute the observed benefits to 'interactive AI features,' but the experiment cannot isolate the contribution of the interactive agent from the automatic cleansing. This is acknowledged in Section 5.3 as a trade-off, but the causal language in the rest of the paper does not temper the claim. The authors should either restrict their conclusions to the overall LAV system or explicitly discuss what a three-condition design (VOV, VOV with automatic cleansing only, and full LAV) would be needed to establish the specific effect of the interactive AI features.","section":"Sections 3.1-3.3 and 5.3"},{"comment":"The Friedman test across four 'conditions' includes 'Overall,' which is not an experimental condition. 'Overall' is a final-questionnaire global rating of the concept of voice-input review creation, not a condition the participants actually experienced in the same way as VOV or LAV. Including this hypothetical rating as a repeated-measures condition inflates the test and makes the reported chi-square (χ2(3) = 20.15) difficult to interpret. The post-hoc comparisons are also only reported for LAV versus Unaided and Overall versus Unaided, omitting the comparisons involving the actual experimental conditions (VOV versus LAV, VOV versus Unaided). The authors should either restrict the statistical test to the two experimental conditions or clearly justify why 'Overall' is a valid within-subjects condition.","section":"Section 4.6"}],"minor_comments":[{"comment":"The key words contain a typo: 'applicaitons' should be 'applications.'","section":"Additional Key Words"},{"comment":"There are several typos: 'administed' should be 'administered,' and 'participantsoverall' is missing a space and should be 'participants' overall.'","section":"Section 3.3.3"},{"comment":"The Spearman correlation of 0.94 is reported without the sample size; since the study has only 14 participants, it would be helpful to state n and show the scatterplot with individual data points (Figure 7). Given the small n, this correlation may be driven by a few participants, so the interpretation should be cautious.","section":"Section 4.4"},{"comment":"The statement 'All of the 20 improvement suggestions were considered helpful' would be clearer if the denominator and the response format are specified; the paper mentions a 'simple thumbs-up mechanism' but does not report how many of the 20 suggestions received thumbs-up versus thumbs-down at the item level.","section":"Section 4.2.3"},{"comment":"The self-efficacy question uses a 0-10 scale while the other Likert items use 1-7; this should be stated explicitly in the text and in Figure 9, since otherwise readers may assume a common scale.","section":"Section 3.3.3"},{"comment":"The prompt listings contain formatting artifacts, including spaces in the middle of variable names (e.g., 'r e f in e I ns t r uc t i on s') and a typo in Appendix C ('genarated'). These should be cleaned up for readability.","section":"Appendix B and Appendix C"},{"comment":"The post-hoc p-values are reported as 'p < 0.001' without the test statistic or the Bonferroni-adjusted significance threshold; providing the actual adjusted p-values and Wilcoxon test statistics would improve reproducibility.","section":"Section 4.6"},{"comment":"In Table 4, the user instruction 'not bricks or balls but maybe crotchets' appears to contain a typo; if the intended word is 'croquettes,' the table entry should be corrected to avoid confusion.","section":"Appendix E"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely topic and the field deployment is a strength, but the empirical support for the headline self-efficacy claim is currently thin because the baseline is retrospective. The willingness-to-share result may survive a corrected analysis, but the confounding of automatic cleansing with the interactive AI agent limits the causal interpretation. I would encourage the editor to request a revision that either adds a pre-intervention baseline, softens the causal claims appropriately, and corrects the statistical reporting, or restricts the claims to what the design can support. The paper could be acceptable after such a revision, but not in its current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this for a quick look at whether voice input plus LLM-assisted editing changes willingness to share restaurant reviews. The short answer: probably yes, based on a small but real field study, but the paper's self-efficacy claim is shakier than it looks because the baseline was measured retrospectively.\n\nThe genuinely useful piece is Table 2—a taxonomy of edit prompts users typed into the AI agent (add details, clarify ambiguity, adjust sentiment, style control). That comes from actual usage and is concrete enough to reuse. The app is an incremental extension of Rambler, but the deployment context (restaurant reviews) and the detailed three-part prompt design in the appendices are properly reported. The authors are transparent about their method.\n\nSoft spots. Most important: the big self-efficacy jump (4.86 to 9.00 on 0–10) compares a retrospective 'Unaided' rating collected in the final questionnaire with post-intervention ratings. The authors acknowledge this in Section 5.3. That means the causal sentence in the abstract—improving users' self-efficacy—is not supported by the data. A retrospective estimate made after using the tool is not a pre-intervention baseline. The willingness-to-share comparison (LAV 5.93 vs VOV 4.53) does not share that problem because it's an in-the-moment rating right after each review, but it still depends on self-report and a small sample (14). Also, the reported t(14) should be t(13); that's minor but sloppy.\n\nWho's this for: HCI researchers studying AI-assisted writing, review platforms, or mobile voice interaction. It's a reasonable conference-level study with a real (if small) deployment, and it does not overclaim in its discussion—the limitations section is honest. The central willingness-to-share finding is plausible and worth taking seriously. The self-efficacy finding should not be cited as evidence of a causal effect.\n\nI'd send this to peer review. A good referee would ask for a revised framing: separate the concurrent sharing result from the retrospective self-efficacy result, and either drop the causal claim or collect a true baseline in a follow-up. The paper deserves that level of engagement, not a desk reject.","headline":"A transparent, small field study; the willingness-to-share effect is defensible, but the self-efficacy claim rests on a retrospective baseline collected after the tool was used, so treat that headline gain as unestablished.","tokens_in":21732,"tokens_out":3088,"would_cite":false,"duration_ms":31850,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A month-long field study with 14 campus diners found that letting a large language model polish a spoken review—while keeping the human's own words, tone, and control—raised users' willingness to share the review and their self-efficacy…","keywords":["online reviews","voice input","large language models","self-efficacy","willingness to share","mobile application","user study","spoken reviews"],"falsifier":"Run the same app with the baseline self-efficacy question asked at onboarding, and track whether participants actually publish their reviews to a live platform instead of just moving a slider. If the onboarding baseline already equals 9/10, or if LAV reviews are not posted more often than VOV reviews, the paper's central claims would not survive.","tokens_in":20799,"feed_emoji":"🎤","tokens_out":10332,"duration_ms":93968,"temperature":0.7,"pith_summary":"Typed reviews are effortful, and people who want to leave detailed feedback while moving through the day often give up before hitting submit. This paper asks whether a spoken review, cleaned up by a large language model, can lower that barrier without replacing the human author. In a counterbalanced field study with 14 adults recruited on a university campus, using a mobile app called Vocalizer, participants reported a significantly higher willingness to share LLM-assisted reviews (mean 5.93 vs 4.53 on a 1–7 scale) and higher self-efficacy—confidence in their own ability to write a good review—after using the assisted version (9.00 vs 4.86 on a 0–10 scale). The authors interpret the result as evidence that AI should edit, not create: users ramble, the model trims and clarifies, and users retain control through direct prompts and optional tips. If correct, this points to a practical way for review platforms to encourage richer user contributions without making people feel their words have been taken over by a machine.","feed_headline":"AI-polished voice reviews boost willingness to share","feed_subtitle":"A 14-person field test finds users more willing to publish AI-refined voice reviews than voice-only drafts.","key_machinery":"The central object is Vocalizer, a mobile web app that converts a spoken restaurant review into a submission-ready text in two versions: voice-only and LLM-assisted. The LLM-assisted version adds three GPT-4-driven features that carry the argument: (1) an initial automatic cleanup that removes filler words, rambling, and off-topic content while keeping the original tone and English level; (2) a conversational AI agent that obeys user-typed instructions to add, omit, correct, clarify, rephrase, or change sentiment in the review; and (3) a suggestion module that generates review-specific tips grounded in published findings on what makes reviews helpful. The design principle is user autonomy: the LLM edits the user's own words rather than generating a review from scratch, and the user can iterate or restart at any point.","core_discovery":"The paper's central claim is that interactive LLM assistance during review writing increases both willingness to share and self-efficacy compared with voice-only transcription, because it converts spontaneous speech into coherent text while preserving the user's content, tone, and control. In the Vocalizer deployment, users recorded up to five minutes of speech, read a polished version, could give free-form instructions to a conversational AI agent, and could request research-grounded improvement tips; all 14 participants preferred the LLM-assisted version. Across 82 LLM-assisted reviews, mean willingness to share was 5.93/7 versus 4.53/7 for 75 voice-only reviews (paired t-test, p < .05), mean satisfaction was 6.15/7, and the AI agent's usefulness averaged 6.09/7. Self-efficacy, defined as confidence in one's ability to give a good review, rose from 4.86/10 unaided to 9.00/10 after the LLM-assisted version, against 8.25 after voice-only, with a significant Friedman test (χ²(3) = 20.15, p < .001) and significant pairwise differences between the unaided baseline and both tool conditions. The authors conclude that the LLM-assisted version increased users' self-efficacy in their review writing abilities.","pith_inferences":["If the willingness-to-share effect is real, the practical payoff is a lower effort barrier for mobile contributions: review platforms could add an 'AI polish' step after voice input and expect both higher output volume and richer content, as users asked for added detail in 30 of the 92 logged prompt modifications.","The paper itself flags in Section 5.3 that the unaided self-efficacy baseline was collected in the final questionnaire, after users had already tried both tools; that ordering likely inflates the reported 4.86-to-9.00 jump, so the effect size is an upper bound.","A three-arm version—voice-only, voice plus automatic cleanup, and voice plus cleanup and interactive agent—would separate the contribution of the conversational agent from the automatic transcription polish, which the paper acknowledges was confounded by design.","The same edit-don't-create pattern could transfer to other user-generated content such as emails, social posts, or accessibility aids for non-native writers, where the relevant question is whether readers perceive the polished text as still authentic; this study measured willingness to share, not reader-side authenticity."],"forward_implications":["Review platforms that add an LLM polish step to voice input can expect higher stated willingness to publish, with the study's effect size (5.93 vs 4.53 on a 7-point scale) and significant paired t-test.","Users will use such agents mostly to add detail and clarify points rather than to fix grammar, so systems should support expansion-prompts well.","A single session with an AI-assisted review tool may raise users' confidence in their review-writing ability, potentially motivating more frequent contributions.","More experienced reviewers write longer, more specific instructions to the agent, implying the tool's value grows with the user's writing skill.","All 14 participants preferred the assisted version, but participant concerns about lost authenticity point to a design tension that any deployment must manage by keeping the author's voice visible."],"supporting_citations":[{"why":"Rambler, the closest prior system, showed that LLM-assisted gist manipulation helps people write from voice; Vocalizer extends this approach to restaurant reviews and adds an interactive edit agent.","marker":"[31]"},{"why":"Empirical grounding for the 'edit, not create' design: AI-generated reviews are perceived as less useful, trustworthy, and authentic when readers know they are AI-generated.","marker":"[4]"},{"why":"The self-efficacy construct and scale guide that the paper adopts to measure users' confidence in writing good reviews.","marker":"[5]"},{"why":"The short user experience questionnaire (UEQ-S) used for the post-visit comparison between the voice-only and LLM-assisted versions.","marker":"[47]"},{"why":"Published findings on what makes an online review helpful, which are fed into the improvement-tips module.","marker":"[36]"},{"why":"The thematic analysis method the authors adapted to code users' AI-agent prompts and open-ended questionnaire responses.","marker":"[7]"}],"fun_headline_variants":["AI-polished voice reviews lift sharing intent","Voice reviews gain AI boost, users share more","LLM-assisted voice reviews raise self-efficacy","AI-enhanced voice reviews increase willingness to post","Interactive AI refines voice reviews, boosts sharing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the self-report measures are truthful: a 1–7 'willingness to share' answer predicts actual sharing, and a single retrospective self-efficacy score, asked after people had already used both tools, reflects their genuine unaided baseline confidence rather than a glow from the tool.","fun_headline_variants_meta":{"raw":{"variants":["AI-polished voice reviews lift sharing intent","Voice reviews gain AI boost, users share more","LLM-assisted voice reviews raise self-efficacy","AI-enhanced voice reviews increase willingness to post","Interactive AI refines voice reviews, boosts sharing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000144,"raw_usage":{"total_tokens":1179,"prompt_tokens":950,"completion_tokens":229,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":566,"completion_tokens_details":{"reasoning_tokens":161}},"tokens_in":566,"tokens_out":229,"duration_ms":2974,"temperature":1.0,"reasoning_tokens":161,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:42:31.240885+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same app with the baseline self-efficacy question asked at onboarding, and track whether participants actually publish their reviews to a live platform instead of just moving a slider. If the onboarding baseline already equals 9/10, or if LAV reviews are not posted more often than VOV reviews, the paper's central claims would not survive.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The self-efficacy construct and scale guide that the paper adopts to measure users' confidence in writing good reviews."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The short user experience questionnaire (UEQ-S) used for the post-visit comparison between the voice-only and LLM-assisted versions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Published findings on what makes an online review helpful, which are fed into the improvement-tips module."}],"review_version":1}