{"id":"a6c3f506-e195-48df-ac87-ce9f66bc23e2","arxiv_id":"2504.14427","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 10-participant Rapid Assessment Process study identifies researcher needs and gaps for SIA development frameworks, applied to the Estuary framework.","lead":"This paper reports a study with 10 SIA researchers to learn what they need from tools that build socially interactive agents. The authors use this feedback to assess their own open-source framework, Estuary, and propose design advice for similar frameworks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Demo-framework conflation in Estuary feedback; conclusions about the framework rely on comments that Section 6.1 admits were often about the demo.","rationale":"The reader's weakest_assumption is exactly the load-bearing concern: the demo is used as a proxy for evaluating the framework. I agree that this is the most serious threat to the central claim, and Section 6.1's own admission strengthens the concern. The issue is not that the case study is worthless, but that the interpretation of RQ3/RQ4 themes as feedback on Estuary's architecture is uncertain. This can be settled by re-coding the existing transcripts to separate demo-anchored from architecture-anchored statements. Because the paper already acknowledges the limitation and frames itself as a case study, the appropriate verdict remains CONDITIONAL; no adjustment to the reader's verdict is needed. If the re-coding were to show systematic contamination, the conclusions would need to be narrowed; if not, the current framing can stand.","tokens_in":12494,"tokens_out":6281,"duration_ms":61011,"concrete_test":"Re-code the 10 existing interview recordings or transcripts (Section 4.3) with two exclusive tags for every Estuary-relevant statement: D=demo-anchored (references the AVP session, AR menu, pathfinding, voice chat, visible latency, or character appearance) or A=architecture-anchored (references Stages, DataPackets, off-cloud/on-prem, client-server, SDK, Python, open-source, or hardware portability). Then compute, per RQ3/RQ4 theme, the fraction of D versus A statements. If the themes cited as evidence that Estuary addresses gaps are predominantly D-anchored, the paper's transfer from demo to framework is unsupported and the conclusions should be reframed as prototype feedback; if predominantly A-anchored, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two parts: the RAP sample yields a trustworthy picture of SIA tool needs, and participants' feedback tells us something about Estuary as a framework. The first part is moderately supported by the protocol (saturation at 10, guided interviews, thematic coding), but the second rests on an assumption the paper itself undermines: that feedback elicited around the AVP demo is feedback about Estuary's architecture. Section 4.3 describes a demo with no NLU, where the planned 'come sit with me' command was replaced by a hand-tracked AR menu, and with ChatGPT 3.5 as a cloud backend. The architecture's distinctive claims (off-cloud execution, modular Stages/DataPackets, Python server, platform-agnostic SDK) were not exercised in the demo; they were only described in the briefing. Interview questions C.3.1-C.3.4 ask for an impression of Estuary immediately after the demo, so the referent is ambiguous. Section 6.1 explicitly admits 'Participants often provided feedback on the demo rather than the framework itself.' The RQ3/RQ4 themes therefore mix demo-level observations (e.g., low-latency voice, AR navigation) with architecture-level judgments. Since the conclusion that researchers perceive Estuary as addressing gaps requires the architecture to be the object of evaluation, this conflation is load-bearing. The paper does not report how many participants were interviewed before versus after the protocol adjustment, so the size of the contamination is unknown.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a case study of user-centered design for Estuary, an open-source multimodal framework for building real-time socially interactive agents. Using the Rapid Assessment Process, the authors conducted one-hour interviews with ten USC ICT-affiliated researchers to (1) identify valued aspects and gaps of current SIA development tools, and (2) evaluate Estuary's design principles. Thematic analysis of interview transcripts yields four sets of themes corresponding to the stated research questions. The paper concludes with design recommendations for SIA frameworks and an Estuary roadmap.","tokens_in":12788,"tokens_out":4512,"duration_ms":38993,"significance":"If the findings are accepted, the paper makes a useful contribution as one of the few published empirical needs analyses of SIA framework users, with a transparent interview guide (Appendix C) and an openly acknowledged set of limitations. The needs themes (RQ1/RQ2), such as integration difficulties, LLM reliability, and sustainability, align with known community concerns and are, in principle, actionable for framework designers. The evaluation of Estuary (RQ3/RQ4) is less secure because it rests on a demo that does not exercise the framework's core architectural claims, and on an in-group sample. The paper is honest in Section 6.1 about the demo–framework conflation and the convenience sample, which strengthens its credibility as a case study, but the framework-level conclusions need to be re-examined.","major_comments":[{"comment":"The RQ3 and RQ4 findings are supposed to evaluate Estuary as a framework, but the evidence was gathered around the AVP demo in which the NLU path was replaced by a hand-tracked AR menu and the conversation was powered by ChatGPT 3.5 in the cloud; Section 6.1 admits that participants often provided feedback on the demo rather than the framework itself. Because interview questions C.3.1–C.3.4 were asked immediately after the demo and do not systematically separate “demo” from “framework” referents, the themes in Sections 5.3–5.4 (e.g., low-latency voice, AR navigation vs. off-cloud, modular Stages/DataPackets, platform agnosticism) conflate two different objects of evaluation. This is load-bearing for the central claim that researchers perceive Estuary as addressing research gaps, and the paper does not report how many interviews used the adjusted protocol that asked for feedback on both.","section":"4.3, 6.1, C.3"},{"comment":"The participants were all recruited from USC ICT within the past five years, and the interviewer was a fellow ICT researcher; this creates an in-group sample for evaluating a framework developed at the same institution. The paper describes the sample as “leading researchers in the field” and generalizes the RQ1–RQ4 findings to the SIA research community, but the convenience sample cannot support that scope of inference. This is acknowledged as a limitation in Section 6.1, but the Discussion still frames the findings as community-level requirements; the claims need to be scaled back to ICT-affiliated researchers, or additional evidence of representativeness is needed.","section":"4.1, 6.1"},{"comment":"The thematic analysis is described as two coders each coding five transcripts with cross-review, but the manuscript does not report how many transcripts were coded by both coders, how disagreements were resolved, or whether the coding was done before or after the interview protocol adjustment. Given the admitted ambiguity between demo and framework feedback, the absence of a per-interview breakdown makes it impossible to assess the contamination size. Please report the protocol adjustment timeline and the coding consensus process.","section":"4.4, 6.1"}],"minor_comments":[{"comment":"The term “leading researchers” overstates the sample; consider using “researchers” or “ICT-affiliated researchers” to match the actual recruitment.","section":"Abstract, Section 1"},{"comment":"The coding description is ambiguous: please clarify whether each transcript was independently coded by two coders or by only one coder with later cross-review, and report inter-coder agreement if applicable.","section":"Section 4.4"},{"comment":"The scoping review is described as systematic, but no search dates, inclusion counts, or PRISMA-style flow are provided; adding this information would support reproducibility.","section":"Section 2"},{"comment":"The protocol adjustment mentioned in Section 6.1 is not documented in the interview guide; please add the adjusted wording and indicate which participants received the updated protocol.","section":"Appendix C.1, Section 6.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is an extended abstract, and its value is mainly as a practice report. The central needs themes may be publishable as-is, but the framework evaluation requires either a reanalysis that separates demo feedback from architecture feedback or a substantial qualification of the RQ3/RQ4 claims. I would not reject the paper, but I would ask for major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hey [Colleague],\n\nRead the Estuary case study. It's a decent, modest paper: the first user-centered design needs assessment for SIA frameworks, at least per their scoping review. That's genuinely new, even if the resulting themes (integration, latency, privacy, open-source) are familiar to anyone in HCI. The RAP protocol, interview guides, and thematic analysis are described in enough detail that someone could replicate the process, and the authors are upfront about limitations. The architecture description of Estuary is also clear.\n\nWhere it's soft: the feedback on Estuary itself is muddled. Participants saw a demo with a cartoon agent on AVP, no NLU (hand menu instead), and ChatGPT 3.5 cloud backend. The authors admit in Section 6.1 that participants often commented on the demo rather than the framework. So the RQ3/RQ4 themes about Estuary's strengths and weaknesses blend demo-level impressions (latency, AR navigation) with architecture-level judgments (off-cloud, modular Stages). They don't report how many interviews happened before the protocol adjustment to separate these, so you can't tell how much contamination exists. That's not fatal for the needs-checklist part, but it substantially weakens the framework evaluation.\n\nAlso, the sample is 10 people from the authors' own institute, interviewed by an insider. RAP calls for insider interviews, so that's a feature, but it still limits generalizability. The paper acknowledges this.\n\nNet: the needs themes and the open methodology are the contribution; the Estuary-specific evaluation is suggestive, not strong. For a CHI EA this is within bounds. I'd send it to review — it's an original case study that gives future framework builders a checklist and a replicable interview protocol. The weaknesses are addressable in follow-up work, not structural.\n\nWorth citing if you're working on SIA frameworks.","headline":"A small, honest RAP case study on SIA framework needs; the demo-framework conflation is real and admitted, but the open methodology and the needs themes make it worth a serious read.","tokens_in":13218,"tokens_out":2581,"would_cite":true,"duration_ms":23021,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Ten researcher interviews chart what SIA frameworks must deliver","keywords":["Socially Intelligent Agents","Conversational AI Framework","Extended Reality","User Testing","Rapid Assessment Process","Multimodal Interaction","User-Centered Design","Framework Evaluation"],"falsifier":"Run the RAP interviews again with a fully functional version of Estuary that includes natural-language understanding and off-cloud LLM support, recruiting participants from multiple institutions; if the recurring themes about Estuary's flexibility and low latency do not reproduce, or if demo-focused comments no longer align with architecture-focused ones, the paper's claim that RAP captures framework-level values and gaps would be weakened.","tokens_in":12336,"feed_emoji":"🤖","tokens_out":3049,"duration_ms":28686,"temperature":0.7,"pith_summary":"This case study argues that Rapid Assessment Process interviews with ten researchers in socially interactive agents (SIAs) produce a dependable picture of what the field values and where current development tools fall short. It further claims that these researchers, after seeing a demo of the open-source multimodal framework Estuary, perceive it as addressing many of those gaps, especially off-cloud operation, platform-agnostic clients, and an interoperable microservice design. If correct, the work shows that a lightweight interview method can steer framework roadmaps and that frameworks built around these priorities will fit real research needs better than current alternatives.","feed_headline":"User interviews chart what SIA frameworks must deliver","feed_subtitle":"Ten researcher interviews turn into a roadmap: open-source, off-cloud, and interoperable tools win.","key_machinery":"The Rapid Assessment Process (RAP) is the methodological engine: directed one-hour insider-to-insider interviews run until thematic saturation around ten participants, with transcripts coded independently by two analysts and grouped into themes. The artifact under evaluation is Estuary's event-based client-server architecture, in which asynchronous 'Stages' wrap microservices in isolated child processes and route multimodal DataPackets over Socket.IO, paired with a Unity SDK for XR clients including the Apple Vision Pro demo used in the study.","core_discovery":"The central claim is that practicing SIA researchers converge on a small set of priorities—open-source availability, off-cloud and on-cloud flexibility, interoperability of microservices, and multimodal support—and that Estuary's design principles align with those priorities. The paper also claims that the main perceived drawbacks of Estuary are the complexity of its client-server setup and the desire for more integrated sensing modalities and controllable dialogue management. These claims are grounded in thematic analysis of one-hour insider interviews, coded by two team members, with themes grouped under four research questions covering current strengths, current gaps, Estuary's perceived value, and its perceived shortcomings.","pith_inferences":["The transferability of the findings is likely limited by the single-institution convenience sample and by the fact that participants commented on an early demo without NLU; a broader sample could shift the priority ordering of the identified themes.","If the perceived value of off-cloud and open-source capabilities holds, competitive pressure on cloud-dependent commercial platforms should increase, and objective benchmarks comparing latency and integration effort across frameworks would become valuable.","A direct testable extension would be to run the same RAP protocol with researchers who have used Estuary in a full study rather than a demo, and check whether the RQ3/RQ4 themes persist.","The desire for standardized message protocols across microservices suggests the field may converge on an interchange standard similar to what Estuary's DataPackets propose."],"forward_implications":["Developers of future SIA frameworks should treat open-source access, off-cloud operation, and microservice interoperability as primary design targets rather than afterthoughts.","Researchers want the ability to run the same framework fully offline or with cloud services interchangeably, so frameworks supporting both modes will fit more study contexts.","The client-server split is a real adoption barrier; providing a standalone one-device mode and simpler networking setup would directly address a top participant concern.","A ten-interview RAP cycle can steer iterative framework roadmaps, making it a viable alternative to large-scale surveys during early design."],"supporting_citations":[{"why":"Supplies the Rapid Assessment Process method that structures the interviews and justifies the ten-participant sample.","marker":"[1]"},{"why":"Prior Estuary publication defining the framework and its design principles under evaluation.","marker":"[10]"},{"why":"Provides the definition of socially interactive agents and the multimodal behavior scope the framework targets.","marker":"[11]"},{"why":"Virtual Human Toolkit serves as a baseline academic framework for comparing current gaps.","marker":"[9]"},{"why":"Pipecat serves as an open-source voice-agent framework for contrast on integration and latency trade-offs.","marker":"[5]"},{"why":"NVIDIA ACE represents a commercial platform illustrating cost, licensing, and engine-lock trade-offs.","marker":"[14]"}],"fun_headline_variants":["Open, flexible, multimodal: SIA framework wishlist","Estuary aligns with SIA priorities but setup is complex","SIA researchers converge on openness, flexibility, multimodality","User interviews map SIA framework priorities and gaps","What SIA developers want: open-source, interoperable, multimodal"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Participants' reactions to the early cartoon demo, which used a hand menu instead of natural-language understanding, are treated as valid evidence about the Estuary framework itself; the authors acknowledge that participants often commented on the demo rather than the architecture, so if the demo misrepresents the framework the conclusions about Estuary's strengths and weaknesses do not transfer.","fun_headline_variants_meta":{"raw":{"variants":["Open, flexible, multimodal: SIA framework wishlist","Estuary aligns with SIA priorities but setup is complex","SIA researchers converge on openness, flexibility, multimodality","User interviews map SIA framework priorities and gaps","What SIA developers want: open-source, interoperable, multimodal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000248,"raw_usage":{"total_tokens":1471,"prompt_tokens":793,"completion_tokens":678,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":409,"completion_tokens_details":{"reasoning_tokens":596}},"tokens_in":409,"tokens_out":678,"duration_ms":6312,"temperature":1.0,"reasoning_tokens":596,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:48:23.638596+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the RAP interviews again with a fully functional version of Estuary that includes natural-language understanding and off-cloud LLM support, recruiting participants from multiple institutions; if the recurring themes about Estuary's flexibility and low latency do not reproduce, or if demo-focused comments no longer align with architecture-focused ones, the paper's claim that RAP captures framework-level values and gaps would be weakened.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Rapid Assessment Process method that structures the interviews and justifies the ten-participant sample."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior Estuary publication defining the framework and its design principles under evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the definition of socially interactive agents and the multimodal behavior scope the framework targets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Virtual Human Toolkit serves as a baseline academic framework for comparing current gaps."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Pipecat serves as an open-source voice-agent framework for contrast on integration and latency trade-offs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"NVIDIA ACE represents a commercial platform illustrating cost, licensing, and engine-lock trade-offs."}],"review_version":1}