REVIEW 3 major objections 5 minor 13 references
HARP: The Human--AI Research Platform
T0 review · 3 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read HARP makes live LLM behavior a controllable experimental variable
desk verdict A clearly written design proposal for a live-LLM experimentation platform; the idea is plausible and useful, but no implementation or pilot data yet, so the central enabling claim is unverified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The configurable live agent: a researcher-facing dashboard sets the agent's system instructions, context, response length, tone, and model parameters for each experimental condition; a participant-facing chat interface embeds the task; and a recording layer captures transcripts along with keystroke-level process measures. The configuration layer is what turns an otherwise unpredictable LLM into something treatable as a set of experimental variables.
What would settle it
Run one configured HARP condition fifty times with different participants and compare the distributions of response content, length, and tone to the same condition run across a model update; if within-condition variation matches or exceeds the between-condition differences planned in the illustrative study, the platform has not yet demonstrated the control it claims.
Extended reading notes
Core claim
The paper's central claim is that HARP places participants in controlled mock scenarios with live, configurable LLM agents, allowing a researcher to hold the task and agent context constant while varying agent prompts, response length, tone, model parameters, and experimental conditions. It pairs this with process-level logging—prompt composition time, response latency, deletions, and keystroke pauses—and triggered in-context surveys. The stated payoff is systematic testing of AI design choices while preserving the dynamic nature of conversational AI, something static prototypes cannot offer and unconstrained live systems cannot guarantee. The paper presents a planned study on technical spec
Load-bearing premise
The load-bearing premise is that a researcher's configuration of prompts, model parameters, and response characteristics is enough to keep the live agent's behavior consistent within a condition and measurably different across conditions, despite stochastic generation and model updates.
Editorial extensions
If this is right
- Researchers can A/B test AI response phrasing, length, and tone using live models rather than static mockups, which the paper argues preserves ecological validity.
- Process measures such as hesitation, deletions, and prompt rewriting make prompt formulation observable instead of relying only on final transcripts.
- Triggered surveys allow in-context feedback at defined points, tying self-report measures to specific moments in the interaction.
- Unmoderated deployment through recruitment platforms or event kiosks allows larger-sample studies without a moderator present.
- Enterprise AI design choices can be treated as experimental variables and evaluated through behavioral, self-report, and task-performance measures before broad rollout.
Reading between the lines
- The platform's control claim is unvalidated until within-condition stability is measured; running the same configured condition repeatedly and comparing response variance across stochastic draws and model versions would be a natural first validation.
- Keystroke measures will likely need per-user baselines because typing style varies widely; the paper itself flags this, suggesting aggregated pause times should be interpreted cautiously even if HARP works.
- The same logging layer could generalize beyond text chat—voice input, facial expression analysis, and gesture recognition are already listed as future work, so HARP may evolve into a multimodal interaction experiment platform.
- If the illustrative study finds retention differences by technicality and length, that would support the broader platform; a null result would not falsify HARP itself, since it would only show that those particular manipulations did not move retention.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces HARP, a proposed web-based platform for human-AI interaction experiments. HARP is designed to let researchers place participants in mock scenarios with live, configurable LLMs, controlling agent prompts, model parameters, response characteristics, and experimental conditions, while logging conversation transcripts, prompt composition time, response latency, deletions, and keystroke pauses. The authors illustrate the platform with a planned study on how response length and technical specificity affect retention of LLM output. The paper is explicitly at the design stage: Section 4 states that data collection has not yet begun. The central claim is that HARP enables systematic testing of AI design choices by treating LLM characteristics as experimental variables while preserving the ecological validity of live conversational AI.
Significance. If the platform works as described, it would address a genuine methodological gap: researchers currently face a trade-off between static, controllable stimuli and unconstrained live LLMs, and they lack process-level logging of how users formulate and revise prompts. The proposed logging dimensions—composition time, deletions, pauses—are a thoughtful complement to transcript analysis. The paper is honest about its current stage, hedges its claims, and explicitly acknowledges one limitation (individual variation in typing behavior). However, the central 'enables' assertion is not yet evidenced: no pilot data, no manipulation check, no comparison with existing tools, and no implementation details are provided. The contribution is currently at the design-proposal level, and its significance will depend on a validation study showing that researcher configuration can indeed yield stable, separable LLM conditions.
major comments (3)
- [Section 4] The central claim that HARP 'enables systematic testing' and 'treat[s] LLM characteristics as experimental variables' (Sections 3 and 5) rests on the assumption that configuring prompts and parameters yields stable, separable conditions. Section 4 explicitly states that 'data collection has not yet begun,' and no pilot data, simulation, or manipulation check is reported. Given that the paper itself cites stochastic generation and model updates as a known difficulty (Section 3, Bommasani et al., 2021), the absence of any evidence that configured conditions produce the intended between-condition differences and within-condition consistency is load-bearing. The authors should provide a validation component, e.g., collect outputs from each configuration and report distributions of length and technicality, or implement a runtime check that rejects or filters outputs outside the intended specs
- [Section 3 / Figure 1] The paper lists 'response length, tone, model parameters, system instructions, and contextual information' as configurable, but it does not specify how these are operationalized within the platform. Is response length controlled through prompt instructions, constrained decoding, or post hoc truncation? Is tone implemented via system prompts, few-shot examples, or model fine-tuning? The illustrative study (Section 4) likewise does not state which model is used, which parameter settings instantiate 'short/medium/long,' or which instructions realize 'developer-oriented' vs. 'layperson-oriented.' Without these details, a reader cannot judge whether the independent variable is actually under researcher control. At minimum, the illustrative study should be specified precisely enough to be implementable.
- [Section 5] The discussion claims that HARP 'reduces reliance on disconnected research tools' and 'offers a way to preserve the responsiveness of conversational AI while improving consistency and experimental control.' These are comparative and performance claims, but the paper does not compare HARP with any existing alternative: static prototypes, Wizard-of-Oz setups, manual prompt-engineering pipelines, or other research platforms. Even a qualitative feature comparison or a small usability test with researchers would help position the contribution. As written, the claims are plausible but not demonstrated.
minor comments (5)
- [Throughout] Several citation formatting issues appear in the text, e.g., 'Amershietal., 2019' and 'Rogers et al., 2011;' with missing spaces, and the reference list has inconsistent spacing ('Shadish,W.R.'). Please normalize citations.
- [Abstract] The abstract uses mismatched quotation marks: 'traditional'' should be 'traditional'. Similar typographic issues appear in the body text.
- [Section 2] The statement 'conversation logs alone provide limited visibility into how users formulate prompts' is asserted without supporting citations. Consider referencing writing-process and keystroke-logging literature, which is already cited later.
- [Sections 2 and 4] The paper describes HARP as 'designed' but does not include any implementation or availability statement. A platform paper should state whether HARP is available as open source, as a hosted service, or as a proprietary tool, and provide at least a high-level architecture description.
- [Section 5] The planned capabilities (voice, facial expression, gesture, emotion analysis) are mentioned without any discussion of privacy, consent, or institutional review beyond 'where legally and ethically appropriate.' A brief paragraph on these issues would strengthen the proposal.
Circularity Check
No circularity: HARP is a platform/design paper with no fitted parameters, no derivation, and no self-referential prediction; its claims are architectural and empirical validation is explicitly deferred.
full rationale
The paper makes no formal derivation or quantitative prediction that could reduce to its own inputs. Its central claim is that a configurable platform can present live LLM agents under researcher-specified prompts, parameters, and response characteristics while logging behavioral measures. That claim is supported by system design (Figures 1-3) and by cited external literature (e.g., Bommasani et al. 2021 for stochasticity; Dhakal et al. 2018 for keystroke variation). No parameter is fitted to data, no 'prediction' is computed from a model built on the same data, and the illustrative study is explicitly described as not yet begun ('Although data collection has not yet begun'), so there are no fitted-input-as-prediction or self-definitional moves. The acknowledged limitations (stochastic generation/model churn, keystroke variability) are validity risks, not circularity. The only self-references are to the authors' own platform design, not to an unverified external theorem used to force a conclusion. Therefore the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Static prototypes and screenshots cannot reproduce the dynamic, open-ended nature of interaction with a live LLM, reducing ecological validity.
- domain assumption Configurable agent prompts, model parameters, and response characteristics yield repeatable experimental conditions with live LLMs.
- domain assumption Keystroke pauses, deletions, and response latency provide insight into cognitive effort or certainty beyond what transcripts show.
invented entities (1)
-
HARP platform
Cite this review
Pith. "Pith review of HARP: The Human--AI Research Platform." pith.science (2026). https://pith.science/paper/5KMKGYPE
@misc{pith2026260720773,
author = {Pith},
title = {Pith review of: HARP: The Human--AI Research Platform},
year = {2026},
howpublished = {\url{https://pith.science/paper/5KMKGYPE}},
note = {Machine review of arXiv:2607.20773}
}
read the original abstract
Large language models (LLMs) have shifted human--computer interaction from `traditional'' interface journeys toward more conversational exchanges. Researchers studying HCI and UI use moderated usability sessions, interviews, surveys, transcript analysis, and static prototypes. However, static prototypes provide limited opportunities to study interaction with live AI systems or systematically control how an LLM behaves across participants and scenarios. Conversation transcripts reveal little about how users formulate, revise, and hesitate over prompts before submission. We designed the Human--AI Research Platform (HARP) for researchers, designers, and anyone who has ever wondered, `What if AI did this?' HARP places participants in controlled mock scenarios with live, configurable AI agents. Researchers can control agent prompts, model parameters, response characteristics, and experimental conditions; trigger surveys at predefined moments; and record prompt composition time, response latency, deletions, and keystroke pauses. Planned capabilities include voice, facial expression, gesture, and, where legally and ethically appropriate, emotion analysis. We illustrate HARP through a study examining how technical specificity and response length affect retention of LLM output. By pairing controllable live agents with behavioral and self-report measures, HARP enables systematic testing of how AI design choices affect users.
Figures
Reference graph
Works this paper leans on
-
[1]
Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems , year =
Dhakal, Vivek and Feit, Anna Maria and Kristensson, Per Ola and Oulasvirta, Antti , title =. Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems , year =
2018
-
[2]
Written Communication , volume=
Keystroke logging in writing research: Using Inputlog to analyze and visualize writing processes , author=. Written Communication , volume=. 2013 , publisher=
2013
-
[3]
, title =
Bucinca, Zana and Malaya, Pooja and Gajos, Krzysztof Z. , title =. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems , year =
2021
-
[4]
arXiv preprint arXiv:2312.10893 , year =
Tankelevitch, Leon and others , title =. arXiv preprint arXiv:2312.10893 , year =
-
[5]
and Dixon, Geoffrey and Bullock, Olivia M
Shulman, Hillary C. and Dixon, Geoffrey and Bullock, Olivia M. and Colón Amill, Diana , title =. Public Understanding of Science , volume =. 2020 , doi =
2020
-
[6]
Keystroke Logging in Writing Research: Using Inputlog to Analyze and Visualize Writing Processes , journal =
Leijten, Mari. Keystroke Logging in Writing Research: Using Inputlog to Analyze and Visualize Writing Processes , journal =
-
[7]
Keystroke Logging in Writing Research: Using Inputlog to Analyze and Visualize Writing Processes , journal =
Leijten, Mari. Keystroke Logging in Writing Research: Using Inputlog to Analyze and Visualize Writing Processes , journal =. 2013 , doi =
2013
-
[8]
and Hursey, Josh and Toscos, Tammy , title =
Rogers, Yvonne and Connelly, Kay and Tedesco, Lucia and Hazlewood, William and Kurtz, Ann and Hall, Robert E. and Hursey, Josh and Toscos, Tammy , title =. UbiComp 2007: Ubiquitous Computing , series =. 2007 , publisher =
2007
Show all 13 references
-
[9]
and Cook, Thomas D
Shadish, William R. and Cook, Thomas D. and Campbell, Donald T. , title =
-
[10]
Rogers, Yvonne and Sharp, Helen and Preece, Jenny , title =
-
[11]
Wizard of Oz Studies -- Why and How , booktitle =
Dahlback, Nils and J. Wizard of Oz Studies -- Why and How , booktitle =
-
[12]
Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems , year =
Amershi, Saleema and Weld, Dan and Vorvoreanu, Mihaela and others , title =. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems , year =
2019
-
[13]
and Adeli, Ehsan and Altman, Russ and Arora, Simran and von Arx, Sydney and Bernstein, Michael S
Bommasani, Rishi and Hudson, Drew A. and Adeli, Ehsan and Altman, Russ and Arora, Simran and von Arx, Sydney and Bernstein, Michael S. and Bohg, Jeannette and Bosselut, Antoine and Brunskill, Emma and others , title =. arXiv preprint arXiv:2108.07258 , year =
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.