{"id":"8971295d-5e5a-4dd5-a042-284c3783ea69","arxiv_id":"2607.20773","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"HARP is a proposed platform for controlled live-LLM experiments, combining configurable AI agents, in-context surveys, and fine-grained keystroke behavior logging.","lead":"HARP is a research platform that lets scientists run controlled experiments with live, configurable AI chat agents while recording how people type, pause, delete, and revise their prompts. It is aimed at researchers who want to test how AI wording, tone, or speed changes user behavior without giving up experimental control.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"HARP's core premise—that configuration yields stable, separable LLM conditions—is untested; without a manipulation check, live stochastic generation may overwhelm the intended contrasts.","rationale":"The reader's conditional verdict and weakest assumption identify exactly the load-bearing issue: the paper claims experimental control over live LLM behavior but provides no evaluation that configuration alone achieves stable, separable conditions. My stress-test sharpens this into a concrete construct-validity problem: the independent variable in the illustrative study is not the researcher's configuration but the actual LLM output, and the platform does not verify that the output matches the intended condition. This concern is internal to the paper's argument—the authors themselves cite stochasticity and model updates (Bommasani et al., 2021) as a known challenge, and Section 4 admits data collection has not begun. I considered whether the lack of code or implementation details should move the verdict to REJECT, but the paper is explicitly a methods proposal with modest claims, and the conditional verdict already imposes the necessary condition: validate the manipulation and release artifacts. Thus I do not alter the reader's verdict. My concrete test is a single pilot manipulation check that would directly resolve whether the central premise holds, and it also tests robustness to model updates, which is the second half of the weakest assumption. No ad hominem is intended; the issue is purely about evidence and internal validity.","tokens_in":3371,"tokens_out":2404,"duration_ms":28612,"concrete_test":"Run a pilot with a fixed model version and the proposed condition configurations (short/medium/long; developer-oriented/layperson-oriented), for at least 20 repeated sessions per condition. Measure generated response length in tokens and a validated technicality/jargon score. Require (a) condition means differ in the intended direction with non-overlapping 95% confidence intervals or a pre-specified effect size, and (b) within-condition variance is small enough not to overwhelm the planned retention effect. Then repeat the pilot after a model update and report whether the manipulation survives. If not, HARP needs a verification or response-rendering layer before it can support the claimed controlled experiments.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on researchers being able to treat LLM characteristics (response length, tone, technicality) as experimental variables. Section 4 explicitly says data collection has not yet begun, so there is no evidence that configuring prompts, model parameters, or response characteristics actually produces the intended between-condition differences or within-condition consistency. This is not merely a missing pilot: live LLMs are stochastic, so the same configuration can yield varying lengths and tones across trials, and model updates change behavior across sessions. If the illustrative study simply instructs the model to be 'short' or 'developer-oriented', the generated outputs may not reliably instantiate these levels, and the condition manipulation becomes a weak or confounded proxy. HARP logs transcripts, latencies, and keystrokes, but logging does not enforce the independent variable. Without a manipulation check—verifying that outputs in each condition actually differ on the intended dimension and are sufficiently homogeneous internally—the proposed experiments lack construct validity. The paper's own cited concern (Bommasani et al., 2021) about stochasticity and model churn is acknowledged but not addressed by any validation step.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces HARP, a proposed web-based platform for human-AI interaction experiments. HARP is designed to let researchers place participants in mock scenarios with live, configurable LLMs, controlling agent prompts, model parameters, response characteristics, and experimental conditions, while logging conversation transcripts, prompt composition time, response latency, deletions, and keystroke pauses. The authors illustrate the platform with a planned study on how response length and technical specificity affect retention of LLM output. The paper is explicitly at the design stage: Section 4 states that data collection has not yet begun. The central claim is that HARP enables systematic testing of AI design choices by treating LLM characteristics as experimental variables while preserving the ecological validity of live conversational AI.","tokens_in":3589,"tokens_out":3683,"duration_ms":34158,"significance":"If the platform works as described, it would address a genuine methodological gap: researchers currently face a trade-off between static, controllable stimuli and unconstrained live LLMs, and they lack process-level logging of how users formulate and revise prompts. The proposed logging dimensions—composition time, deletions, pauses—are a thoughtful complement to transcript analysis. The paper is honest about its current stage, hedges its claims, and explicitly acknowledges one limitation (individual variation in typing behavior). However, the central 'enables' assertion is not yet evidenced: no pilot data, no manipulation check, no comparison with existing tools, and no implementation details are provided. The contribution is currently at the design-proposal level, and its significance will depend on a validation study showing that researcher configuration can indeed yield stable, separable LLM conditions.","major_comments":[{"comment":"The central claim that HARP 'enables systematic testing' and 'treat[s] LLM characteristics as experimental variables' (Sections 3 and 5) rests on the assumption that configuring prompts and parameters yields stable, separable conditions. Section 4 explicitly states that 'data collection has not yet begun,' and no pilot data, simulation, or manipulation check is reported. Given that the paper itself cites stochastic generation and model updates as a known difficulty (Section 3, Bommasani et al., 2021), the absence of any evidence that configured conditions produce the intended between-condition differences and within-condition consistency is load-bearing. The authors should provide a validation component, e.g., collect outputs from each configuration and report distributions of length and technicality, or implement a runtime check that rejects or filters outputs outside the intended specs","section":"Section 4"},{"comment":"The paper lists 'response length, tone, model parameters, system instructions, and contextual information' as configurable, but it does not specify how these are operationalized within the platform. Is response length controlled through prompt instructions, constrained decoding, or post hoc truncation? Is tone implemented via system prompts, few-shot examples, or model fine-tuning? The illustrative study (Section 4) likewise does not state which model is used, which parameter settings instantiate 'short/medium/long,' or which instructions realize 'developer-oriented' vs. 'layperson-oriented.' Without these details, a reader cannot judge whether the independent variable is actually under researcher control. At minimum, the illustrative study should be specified precisely enough to be implementable.","section":"Section 3 / Figure 1"},{"comment":"The discussion claims that HARP 'reduces reliance on disconnected research tools' and 'offers a way to preserve the responsiveness of conversational AI while improving consistency and experimental control.' These are comparative and performance claims, but the paper does not compare HARP with any existing alternative: static prototypes, Wizard-of-Oz setups, manual prompt-engineering pipelines, or other research platforms. Even a qualitative feature comparison or a small usability test with researchers would help position the contribution. As written, the claims are plausible but not demonstrated.","section":"Section 5"}],"minor_comments":[{"comment":"Several citation formatting issues appear in the text, e.g., 'Amershietal., 2019' and 'Rogers et al., 2011;' with missing spaces, and the reference list has inconsistent spacing ('Shadish,W.R.'). Please normalize citations.","section":"Throughout"},{"comment":"The abstract uses mismatched quotation marks: 'traditional'' should be 'traditional'. Similar typographic issues appear in the body text.","section":"Abstract"},{"comment":"The statement 'conversation logs alone provide limited visibility into how users formulate prompts' is asserted without supporting citations. Consider referencing writing-process and keystroke-logging literature, which is already cited later.","section":"Section 2"},{"comment":"The paper describes HARP as 'designed' but does not include any implementation or availability statement. A platform paper should state whether HARP is available as open source, as a hosted service, or as a proprietary tool, and provide at least a high-level architecture description.","section":"Sections 2 and 4"},{"comment":"The planned capabilities (voice, facial expression, gesture, emotion analysis) are mentioned without any discussion of privacy, consent, or institutional review beyond 'where legally and ethically appropriate.' A brief paragraph on these issues would strengthen the proposal.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"This is a design proposal with a useful idea but no validation. The missing manipulation check is the key issue: as the stress-test note rightly observes, configuring an LLM is not the same as controlling it. The authors should be asked to add a validation study or to reframe the paper as a design exploration rather than claiming experimental enablement. The paper is not appropriate for acceptance in its current form, but the gap is fixable with a pilot/manipulation-check section."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is a methods proposal, not a study. HARP isn't built yet and the paper says so explicitly—Section 4 states data collection has not begun. Read that way, it's a clear and honest design note. The novelty is the combination: live configurable LLM agents, in-context surveys, and keystroke-process logging. That's a plausible and useful integration for HCI researchers who want more ecological validity than static prototypes without losing all control.\n\nWhat the paper does well: the writing is clear, the illustrative study (response length and technicality on retention) is a sensible vehicle, and the authors explicitly flag the main weakness of keystroke measures—typing variability. They also cite the known LLM stochasticity/churn problem rather than ignoring it.\n\nThe soft spots are real but not fatal. The central claim is that HARP 'enables' controlled experiments with live LLMs, but there's no evidence that configuring prompts and model parameters actually yields stable, separable conditions. Without a manipulation check, live generation could wash out the intended contrasts. That's a gap, not a hidden flaw, because the paper is honest about its current state. The more significant omission is the lack of any survey of existing platforms; without that, the novelty claim is unverified.\n\nI'd send this to peer review, but with the clear expectation that the authors add either a pilot or a feasibility demonstration. As a standalone contribution it's a well-scoped proposal, not a validated system. It deserves referee time precisely because the idea is plausible and the field needs better tooling.","headline":"A clearly written design proposal for a live-LLM experimentation platform; the idea is plausible and useful, but no implementation or pilot data yet, so the central enabling claim is unverified.","tokens_in":4007,"tokens_out":1776,"would_cite":false,"duration_ms":19819,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HARP makes live LLM behavior a controllable experimental variable","keywords":["Human-AI interaction","configurable AI agents","experimental control","keystroke logging","prompt composition","response latency","conversational AI","research platform"],"falsifier":"Run one configured HARP condition fifty times with different participants and compare the distributions of response content, length, and tone to the same condition run across a model update; if within-condition variation matches or exceeds the between-condition differences planned in the illustrative study, the platform has not yet demonstrated the control it claims.","tokens_in":3283,"feed_emoji":"🤖","tokens_out":3299,"duration_ms":28502,"temperature":0.7,"pith_summary":"The paper proposes HARP, a research platform that places participants in controlled mock scenarios with live, configurable LLM agents. Researchers can set the agent's prompt, response length, tone, context, and model parameters, then trigger surveys at predefined moments while the platform records conversation transcripts, prompt composition time, response latency, deletions, and keystroke pauses. The central claim is that this pairing closes the gap between static prototypes, which are controllable but artificial, and unconstrained live systems, which are realistic but hard to control. This would enable systematic A/B testing of AI design choices with the responsiveness of real conversational AI. The paper illustrates the approach with a planned study on how response length and technical language affect retention, but explicitly notes that data collection has not yet begun.","feed_headline":"Live LLM behavior becomes an experimental variable","feed_subtitle":"HARP logs pauses, revisions, and latency while researchers vary prompts and model settings.","key_machinery":"The configurable live agent: a researcher-facing dashboard sets the agent's system instructions, context, response length, tone, and model parameters for each experimental condition; a participant-facing chat interface embeds the task; and a recording layer captures transcripts along with keystroke-level process measures. The configuration layer is what turns an otherwise unpredictable LLM into something treatable as a set of experimental variables.","core_discovery":"The paper's central claim is that HARP places participants in controlled mock scenarios with live, configurable LLM agents, allowing a researcher to hold the task and agent context constant while varying agent prompts, response length, tone, model parameters, and experimental conditions. It pairs this with process-level logging—prompt composition time, response latency, deletions, and keystroke pauses—and triggered in-context surveys. The stated payoff is systematic testing of AI design choices while preserving the dynamic nature of conversational AI, something static prototypes cannot offer and unconstrained live systems cannot guarantee. The paper presents a planned study on technical spec","pith_inferences":["The platform's control claim is unvalidated until within-condition stability is measured; running the same configured condition repeatedly and comparing response variance across stochastic draws and model versions would be a natural first validation.","Keystroke measures will likely need per-user baselines because typing style varies widely; the paper itself flags this, suggesting aggregated pause times should be interpreted cautiously even if HARP works.","The same logging layer could generalize beyond text chat—voice input, facial expression analysis, and gesture recognition are already listed as future work, so HARP may evolve into a multimodal interaction experiment platform.","If the illustrative study finds retention differences by technicality and length, that would support the broader platform; a null result would not falsify HARP itself, since it would only show that those particular manipulations did not move retention."],"forward_implications":["Researchers can A/B test AI response phrasing, length, and tone using live models rather than static mockups, which the paper argues preserves ecological validity.","Process measures such as hesitation, deletions, and prompt rewriting make prompt formulation observable instead of relying only on final transcripts.","Triggered surveys allow in-context feedback at defined points, tying self-report measures to specific moments in the interaction.","Unmoderated deployment through recruitment platforms or event kiosks allows larger-sample studies without a moderator present.","Enterprise AI design choices can be treated as experimental variables and evaluated through behavioral, self-report, and task-performance measures before broad rollout."],"fun_headline_variants":["Control live LLMs, record every hesitation","HARP: a platform to tune AI and measure users","Live AI agents, editable parameters, detailed logs","Set AI behavior, then watch users react and revise"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that a researcher's configuration of prompts, model parameters, and response characteristics is enough to keep the live agent's behavior consistent within a condition and measurably different across conditions, despite stochastic generation and model updates.","fun_headline_variants_meta":{"raw":{"variants":["Control live LLMs, record every hesitation","HARP: a platform to tune AI and measure users","Live AI agents, editable parameters, detailed logs","Set AI behavior, then watch users react and revise"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000261,"raw_usage":{"total_tokens":1414,"prompt_tokens":713,"completion_tokens":701,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":457,"completion_tokens_details":{"reasoning_tokens":652}},"tokens_in":457,"tokens_out":701,"duration_ms":7264,"temperature":1.0,"reasoning_tokens":652,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T09:23:01.493157+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run one configured HARP condition fifty times with different participants and compare the distributions of response content, length, and tone to the same condition run across a model update; if within-condition variation matches or exceeds the between-condition differences planned in the illustrative study, the platform has not yet demonstrated the control it claims.","supporting_citations":[],"review_version":1}