{"id":"29f264b9-ba6f-421d-8902-a38e751999a4","arxiv_id":"2508.12388","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Users perceive virtual agents as more trustworthy and motivating when agents display visible, plausible effort as co-participants rather than simply delivering encouragement.","lead":"This pilot study used interviews with 12 people to ask whether virtual fitness agents work better when they seem to be putting in effort alongside the user, not just giving orders. Most people in the study distrusted agents unless their activity looked believable and showed genuine effort.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim rests solely on retrospective self-report from 12 participants; no behavioral or independent-rater evidence is reported, so the perceived-effort mechanism remains unvalidated.","rationale":"The reader's verdict is UNVERDICTED because the supplied full text is garbled, so only the abstract can be evaluated. My stress-test agrees: the decisive question is not whether the qualitative observation is interesting, but whether it supports the causal-sounding claim that perceived genuine effort is a key determinant of trust/motivation. The abstract's evidence is a thematic analysis of 12 retrospective interviews from a prior intervention. That design cannot distinguish between a real psychological mechanism and the participants' (or coders') post-hoc narrative. No behavioral measure—e.g., subsequent engagement, step counts, or validated trust scales—is reported in the abstract. The garbled full text includes fragments of unrelated statistical content and a different arXiv ID, making independent verification impossible. Thus the central claim remains plausible but unsupported. The appropriate verdict is unchanged: the paper should remain UNVERDICTED pending access to the clean manuscript and ideally a confirmatory behavioral experiment. I credit the authors for framing the work as a pilot and for proposing concrete design directions; those are appropriately tentative. However, the 'key determinant' language in the abstract goes beyond what a small qualitative pilot can establish. No formal verification or reproducible code is relevant here. The check I recommend would settle whether the qualitative tension has downstream behavioral consequences.","tokens_in":9552,"tokens_out":4203,"duration_ms":49590,"concrete_test":"Obtain the clean full text and verify the Methods; then run a preregistered between-subjects experiment (e.g., N≈120/arm) in which an agent's displayed step counts are plausible and effortful vs implausible/static, measuring trust and actual step count/app engagement. If the plausible-effort condition does not significantly improve trust, motivation, or behavior, the qualitative tension does not translate into the claimed determinant. This directly tests the causal mechanism behind the central claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that a virtual agent's perceived genuine effort—anchored in believable, relatable activity—drives user trust and motivation. The load-bearing premise is that retrospective interviews with 12 participants from a prior intervention give an unbiased view of their actual psychological responses, and that the thematic analysis did not impose the reported performance–authenticity tension. The abstract reports neither the interview protocol, the coding scheme, inter-rater reliability, nor member checks; it also reports no behavioral outcome. The supplied full text is unreadable and even contains a fragment of a different arXiv preprint ('arXiv:2508.12380v1 [math.ST]'), so the Methods cannot be independently checked. In this situation, 'ambiguous or implausible activity levels undermined trust and motivation' could reflect demand characteristics, recall bias, or researcher-imposed coding rather than a stable mechanism. The design directions (behavioral cues, narrative grounding, personalized performance) are only justified if that qualitative tension reliably predicts engagement. This is not an internal inconsistency; the paper is explicitly a pilot and exploratory. It is an evidentiary gap: the claim may be true, but the present evidence does not yet establish it beyond hypothesis generation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a pilot qualitative study on virtual agents as co-participants in physical activity interventions. Based on thematic analysis of semi-structured interviews with 12 participants from a prior intervention, it claims a recurring tension between perceived performance and authenticity: users value social features when agents or others appear to be genuinely trying, while ambiguous or implausible activity levels reduce trust and motivation. The paper then proposes early design directions for fostering co-experienced exertion, including behavioral cues, narrative grounding, and personalized performance. Only the abstract is readable in the provided source; the full text is corrupt/unreadable and even contains a fragment of an unrelated arXiv preprint (arXiv:2508.12380v1 [math.ST]).","tokens_in":9846,"tokens_out":2849,"duration_ms":35219,"significance":"If the qualitative finding is substantiated, it offers a useful and falsifiable design hypothesis for socially resonant virtual agents: perceived genuine effort, anchored in believable and relatable activity, may be a key determinant of trust and motivation. The paper is explicitly exploratory, and the proposed design directions are plausible extensions of the reported participant accounts. However, the current evidentiary base is thin—retrospective self-reports from 12 participants, with no visible behavioral measures, comparison conditions, coding reliability metrics, or member checks. As it stands, the contribution is best characterized as hypothesis generation rather than validated mechanism. The paper's strength is that it articulates a concrete design-relevant tension that can be tested in future intervention studies; its weakness is that the present evidence is not sufficient to establish that tension as a stable psychological mechanism.","major_comments":[{"comment":"The full manuscript text is unreadable in the provided source: it appears as encoding corruption and contains an unrelated fragment of arXiv:2508.12380v1 [math.ST]. Consequently, the interview protocol, coding scheme, analysis steps, participant demographics, and supporting participant quotes cannot be inspected. The central claim rests entirely on this thematic analysis. Please supply a clean, readable manuscript and verify file integrity.","section":"Full text (provided source)"},{"comment":"The central claim (a performance–authenticity tension) is supported only by retrospective self-reports from 12 participants in a prior intervention. The abstract reports no inter-rater reliability, member checks, negative-case analysis, or triangulation with behavioral logs or system data. Without such evidence, 'ambiguous or implausible activity levels undermined trust and motivation' could reflect recall bias or demand characteristics rather than a stable mechanism. Provide coding reliability metrics and evidence that the thematic coding did not impose the reported tension.","section":"Abstract"},{"comment":"The sample and the prior intervention are not described: What was the prior intervention? What were the eligibility criteria for the 12 participants? How was the interview guide developed? Was thematic saturation reached? Additionally, the phrase 'ambiguous or implausible activity levels' is not operationalized. Please specify how plausibility was assessed and report how many participants expressed each theme, so readers can judge prevalence versus idiosyncratic views.","section":"Abstract"},{"comment":"The paper moves from a qualitative tension to design directions (behavioral cues, narrative grounding, personalized performance) without evaluating those directions. This is an inductive leap, not a validated implication. The proposed directions should be explicitly labeled as hypotheses requiring user evaluation, and any preliminary design rationale should be separated from the empirical findings.","section":"Abstract / Design directions"}],"minor_comments":[{"comment":"The terms 'co-participant' and 'co-experienced exertion' are used without operational definitions; define them early and clarify how they differ from existing constructs such as social presence or companionship.","section":"Abstract / Introduction"},{"comment":"The notion of 'relatable human benchmarks' is vague. Provide concrete examples from participant accounts and explain how benchmarks were elicited or inferred.","section":"Abstract / Findings"},{"comment":"The embedded fragment of arXiv:2508.12380v1 [math.ST] is a serious file-integrity issue. Remove it and confirm the submitted PDF matches the intended manuscript.","section":"Full text"},{"comment":"The abstract and visible fragments do not mention limitations such as small sample size, retrospective nature of the data, or lack of behavioral validation. Add an explicit limitations paragraph if not already present.","section":"Limitations (if present)"}],"recommendation":"major_revision","confidential_remarks":"I cannot verify the central claims because the supplied full text is unreadable and contaminated with an unrelated preprint fragment. This may be a submission/pipeline issue, but under the present circumstances the manuscript does not meet the standard for acceptance. I am not recommending rejection because the abstract is coherent and the claims are plausible exploratory findings; however, the revision must provide a clean full text and substantially strengthen the methodological transparency (coding reliability, participant recruitment, analysis protocol) before the performance–authenticity tension can be considered established."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea here is actually worth a look: reframing virtual agents from coaches into co-participants who visibly exert effort alongside the user. That's a modest but genuine shift, and the specific finding — that implausible or ambiguous activity levels undermine trust and motivation — is concrete and actionable. The paper is honest about being a pilot, and the design directions (behavioral cues, narrative grounding, personalized performance) follow sensibly from the interviews.\n\nThat said, the evidentiary support is as thin as the abstract suggests. Twelve retrospective interviews from a prior intervention, no coding reliability, no member checks, no behavioral outcomes. The tension between performance and authenticity could easily reflect demand characteristics or the researchers' own framing. The reader's take gives this a low confidence and I agree: the mechanism is plausible, not validated.\n\nOne more thing I have to flag: the full text I was given is garbled and actually contains a fragment from a different arXiv preprint (math.ST). I could not check the interview protocol, coding scheme, or analysis. That might be a pipeline artifact rather than an author problem, but it means I can't verify the methods at all. For the record, the abstract and the conclusions it states are internally coherent, so this does not look like a case of overclaiming — it's a scoped exploratory study.\n\nThe audience is HCI people working on virtual coaches and behavior-change apps. A serious referee could push for more rigor (independent coders, a clearer link between the qualitative themes and the proposed design directions, maybe a small follow-up with behavioral measures), but the paper deserves that referee time rather than a desk rejection. The framing is new enough and the topic is practically useful. I wouldn't cite it in my own work yet, but it's a reasonable reading-group pick for a conversation about qualitative evidence in agent design.\n\nRecommendation: send it to review, ask for revisions that strengthen the evidentiary transparency, and treat it as a hypothesis-generating pilot.","headline":"A plausible reframing of virtual agents as co-participants, but the evidence is thin and the methods are unverifiable from the supplied text; still worth referee time as a pilot study.","tokens_in":10242,"tokens_out":2128,"would_cite":false,"duration_ms":28114,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A virtual exercise agent motivates users not through its message but through whether its displayed activity looks like genuine effort.","keywords":["virtual agents","physical activity interventions","co-participant","perceived effort","authenticity","trust","social comparison","behavior change"],"falsifier":"Run a controlled trial with two versions of the same coach agent giving identical messages, one with plausible, effortful activity (realistic step counts, pauses, fatigue cues) and one with perfect, unexplained high activity, and compare user trust and motivation over several weeks. If the two conditions produce the same outcomes, the claim that visible genuine effort drives motivation fails.","tokens_in":9509,"feed_emoji":"🏃","tokens_out":4885,"duration_ms":55588,"temperature":0.7,"pith_summary":"The paper asks whether a virtual agent can motivate physical activity by acting as a co-participant who visibly works alongside the user, rather than only as an instructor who delivers encouragement. Based on interviews with 12 participants from an earlier activity intervention, it reports that users treat the agent's activity levels as evidence of whether it is genuinely trying. When the agent's performance is plausible, effortful, and tied to relatable human benchmarks, users trust it and respond to it; when activity levels are ambiguous, implausible, or unexplained, trust and motivation fall. The authors conclude that perceived authenticity of effort may matter more than message content, and they offer early design directions for making an agent's exertion visible and believable.","feed_headline":"Fitness agents need believable effort to keep users motivated","feed_subtitle":"Twelve users said plausible, effortful activity builds trust; implausible numbers kill motivation.","key_machinery":"Perceived genuine effort in a social comparison context is the mechanism doing the work. The paper treats the believability of the agent's reported activity, and the visible signs of exertion attached to it, as the bridge between the agent's output and the user's trust and motivation. Its proposed design levers—behavioral cues, narrative grounding, and personalized performance—are all ways of making that perceived effort credible.","core_discovery":"The study's central claim is that people evaluate a virtual exercise agent by asking whether its reported performance reflects real effort, and this judgment determines whether the agent can motivate them. In social comparison, users do not just compare numbers; they assess whether the other party is really trying. Participants valued social features when they believed others, including the agent, were genuinely making an effort, and they became skeptical when activity levels seemed ambiguous or too good to be true. The paper names the resulting design problem a tension between perceived performance and authenticity, and argues that agents should be designed to be seen as exerting themselves","pith_inferences":["[Editorial inference] The design directions are not yet tested; a direct experiment varying only the agent's visible effort cues would show whether perceived effort is truly the active ingredient.","[Editorial inference] If perceived authenticity is what drives motivation, then evaluations of such agents should measure trust and perceived genuineness, not just step counts or message recall.","[Editorial inference] The logic likely extends beyond exercise to any persuasive agent that reports performance—diet, sleep, or learning—where implausible reports may trigger the same skepticism.","[Editorial inference] Engineering visible effort raises an ethical question the paper leaves open: making an agent look like it is struggling is manufactured authenticity, and designers would need to decide whether that is honest representation or persuasion by illusion."],"forward_implications":["Physical activity agents should surface effort and struggle, not only achievements, because visible exertion is what makes performance believable.","Agents that display ambiguous or implausible activity levels are likely to reduce trust and motivation, so performance numbers need to be explicable to users.","Grounding agent activity in relatable human benchmarks could make social comparison features more motivating.","Personalizing the agent's performance to the user's own level may preserve believability and prevent the agent from seeming unrelatable.","The co-participant role, rather than the coach role, is a viable design stance for virtual agents in behavior-change interventions."],"supporting_citations":[],"fun_headline_variants":["Effort beats message in fitness agent motivation","Virtual coaches should sweat, not just cheer","Users trust fitness agents that show real effort","Implausible activity kills fitness agent motivation"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The conclusion rests on twelve people's retrospective interview accounts being a faithful and unbiased record of how they actually responded to the agent, and on the coding of those interviews not imposing the reported performance–authenticity tension.","fun_headline_variants_meta":{"raw":{"variants":["Effort beats message in fitness agent motivation","Virtual coaches should sweat, not just cheer","Users trust fitness agents that show real effort","Implausible activity kills fitness agent motivation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1120,"prompt_tokens":689,"completion_tokens":431,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":433,"completion_tokens_details":{"reasoning_tokens":388}},"tokens_in":433,"tokens_out":431,"duration_ms":5458,"temperature":1.0,"reasoning_tokens":388,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:28:16.948663+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled trial with two versions of the same coach agent giving identical messages, one with plausible, effortful activity (realistic step counts, pauses, fatigue cues) and one with perfect, unexplained high activity, and compare user trust and motivation over several weeks. If the two conditions produce the same outcomes, the claim that visible genuine effort drives motivation fails.","supporting_citations":[],"review_version":1}