{"id":"ee3b2062-0ba4-455c-8a55-f2177df6dd00","arxiv_id":"2506.10079","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Audience members felt they controlled a dancer-robot duet by voting, but consistent vote patterns across four shows suggest the system channeled their choices more than it empowered them.","lead":"This paper reports on Dance Squared, a live performance where audience members voted on phones to steer a small robot crawling across a dancer's costume. The authors found that viewers felt they shaped the show, but voting patterns were similar across four nights, suggesting the design guided their choices rather than handing them real power.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Under the unlimited-voting mechanic (§5.2), override ratios are vote-share, not participant-share; without unique-voter counts, the 'strikingly consistent' patterns in Table 2 could be produced by a few hyper-voters, so the claim that collective choices were subtly steered lacks its empirical base.","rationale":"I read the paper as a performance-led design inquiry whose contribution is the artifact and the conceptual distinction between felt agency, agentive behavior, and actual power. That contribution does not require a controlled experiment, and the authors are explicit that the findings are situated and provisional. However, the abstract and conclusion assert a factual pattern — 'voting data ... strikingly consistent patterns' and 'collective behavior ... followed consistent patterns' — and those assertions rest entirely on Figure 8 and Table 2. The unlimited-vote mechanic, deliberately kept after the §4.3 incident, makes vote totals a poor proxy for the number of people choosing. The reported normalization cannot repair this, and no operator log verifies that the winning option changed robot behavior. Thus the empirical demonstration of the headline tension is weaker than the prose suggests. The proposed reprise with tokenized unique-voter counts and a command log would settle whether the pattern is collective and whether the votes affected the robot. Given the authors' own framing as provocation rather than proof, this is a fixable evidentiary gap, not a fatal flaw; the reader's CONDITIONAL verdict remains appropriate.","tokens_in":18130,"tokens_out":5879,"duration_ms":78790,"concrete_test":"Conduct a fifth performance (or a controlled reprise) with the same script but issue each joining device an anonymous random session token at the QR-code landing page; log one token per vote without storing any other identifier. Then recompute the Table 2 override ratios using only the first vote per token per prompt, and separately compute unique-voter turnout per prompt. If the one-vote-per-person ratios deviate from the reported mu by more than binomial sampling error, or if unique-voter turnout falls below 50% of attendees at any prompt, the reported consistency cannot be attributed to collective decision-making. In the same run, have the backstage dashboard append a timestamped robot-command log and verify that the stated winning option was actually executed during each prompt window.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.2 states that 'participants were not required to vote, and were allowed to vote multiple times per prompt'; the reported mu and sigma/mu in Table 2 are normalized ratios of override votes to total votes at each prompt. A ratio of votes is not a ratio of people. If, say, five engaged audience members each cast dozens of votes while 150 others cast one, the override fraction is determined by those five, and the a-f arc in Figure 8 describes their behavior, not a collective decision. The normalization ('to account for variability in total votes') only rescales each prompt to a common total; it cannot recover the missing denominator of unique voters. Consequently the paper's central empirical claim — that 'their collective behavior across four performances followed consistent patterns' — is not supported by the logged data as described. The consistency (sigma/mu 0.034–0.135) could equally reflect a stable propensity of a small hyper-voting subset, or even the timing and placement of prompts. The absence of an operator-side robot-command log is a second, compounding gap: the survey item 'My choices affected the robot's behavior' may track the projected vote visualization rather than any verified change in robot motion, so the 'actual power' half of the felt-vs-exercised distinction is also unmeasured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Dance^2, a 15-minute interactive dance performance in which audience members vote via their phones to either continue or override the choreography of a wearable robot attached to a dancer. The authors describe the five-part performance structure, the technical implementation based on the Calico on-body robot, a ten-month performance-led design process, and results from four public performances. Reported findings are two-fold: a post-performance survey (150 voluntary respondents) suggests that audience members felt a strong sense of connection and perceived influence, while logged voting data from six prompts in Part 4 show override ratios that the authors describe as 'strikingly consistent' across performances (Table 2, σ/μ between 0.034 and 0.135). The paper interprets this tension between felt agency and consistent voting outcomes as evidence that the audience's collective choices were subtly shaped by choreography, framing, and interface design, drawing on Breel's distinction between agentive behavior and the experience of agency and on broader HCI discourse about perceived versus actual control.","tokens_in":18349,"tokens_out":3207,"duration_ms":44625,"significance":"The paper's strength is its honest, situated account of a designed interactive performance and its explicit engagement with a theoretically meaningful distinction—agentive behavior, experience of agency, and actual power. The authors are transparent about the exploratory nature of the work, explicitly disclaiming generalizability, and they avoid fitted parameters or over-quantified claims about motivation. The artifact itself is a reasonable case study for performance-led HCI research, and the discussion of system design, choreographic framing, and ethical data collection contains useful provocations. If the empirical claims about the consistency of collective voting and the felt-versus-exercised power gap can be adequately supported, the paper would make a valuable contribution to HCI discourse on agency in participatory and algorithmically curated systems. However, as presented, the central empirical conclusion rests on a methodological premise that the data do not substantiate.","major_comments":[{"comment":"The voting data are logged as vote counts, not participant counts, because 'participants were not required to vote—and were allowed to vote multiple times per prompt' (§5.2). The normalization described in §5.2 rescales each prompt's total votes to a common scale; it cannot recover the number of unique voters per prompt. Consequently, the override ratios in Table 2 are vote shares, not person shares. If a small subset of highly engaged audience members cast many votes at some prompts, the reported μ and σ/μ would describe that subset's behavior rather than a collective decision. The paper's central claim—that 'their collective behavior across four performances followed consistent patterns' (§7) and that this consistency demonstrates 'subtle shaping' by system design—is therefore not supported by the logged data as described. The authors could remedy this by reporting per-prompt unique-voter distributions, or, if such data are unavailable, by explicitly reframing the finding as vote-level consistency with the caveat that it may reflect hyper-voting dynamics.","section":"§5.2, Table 2, Figure 8"},{"comment":"The survey findings are reported as percentages without per-item response counts, confidence intervals, or any adjustment for the fact that only 150 of over 200 audience members responded voluntarily. For example, '81% of respondents agreed' and '64–65% felt emotionally or perceptually connected' are stated without the denominator for each item (e.g., whether all 150 answered every item) and without any measure of uncertainty. This makes it difficult to assess the stability of the perceived-agency claim, and it also precludes a meaningful comparison between the survey results and the voting data. The authors should report per-item n and, where possible, confidence intervals or at least note item-level missingness; they should also acknowledge that voluntary response may overrepresent engaged audience members, which is relevant because the paper's argument contrasts felt agency with actual collective behavior.","section":"§5.1, Figure 7"},{"comment":"The paper never verifies that the audience's 'winning' choice actually changed the robot's behavior. The system description states that a backstage operator controls the robot via a dashboard, and the audience interface streams voting data, but no operator-side robot-command log or time-synced log of robot state is presented. Without such a log, the survey item 'My choices affected the robot's behavior' (Figure 7) can only be interpreted as a report about perceived influence, possibly driven by the projected vote visualization and the dancer's reactions, not about verifiable changes in robot motion. This means the 'actual power' half of the felt-versus-exercised-power distinction is unmeasured. The authors should provide a log of robot commands or otherwise explicitly state that no behavioral verification exists and adjust the discussion accordingly.","section":"§3.1, §5.2"},{"comment":"The 'strikingly consistent' description is based only on descriptive statistics (μ and σ/μ) computed across four performances, with no baseline, null model, or inferential test. With four data points per prompt, the observed variability could be consistent with chance or with stable prompt-specific properties (such as the action being offered) that have nothing to do with collective agency. Moreover, the prompts differ in action and dramatic context, so the a–f arc could reflect the wording of each choice rather than an emergent collective decision. The authors should provide a more principled comparison (e.g., a permutation test with a null model of random voting) or, at minimum, temper the wording from 'strikingly consistent' to a more qualified claim about apparent similarity across shows.","section":"§5.2, Table 2"}],"minor_comments":[{"comment":"The conclusion contains a typo: 'four performances' is written as 'fours performances'.","section":"§7"},{"comment":"The caption states that the charts show 'a normalized ratio between votes to continue the choreography and votes to override the choreography,' while Table 2 reports the override ratio; the wording should make clear that the plotted quantity is the override share (override votes divided by total votes), not a ratio of two vote counts.","section":"Figure 8 caption"},{"comment":"The labels 'avg votes' and 'stdev' are not defined in the caption or text; it would help to state that these are the mean and standard deviation of total votes per prompt across the performances, or clarify what they refer to if that is not the case.","section":"Figure 8"},{"comment":"The bar chart in Figure 7 would be more informative if the exact percentage and number of respondents for each Likert item were provided, rather than only a visual distribution.","section":"§5.1"},{"comment":"The statement that 'Had we tracked individual behaviors—vote timing, frequency, shifts in response—we might have constructed more detailed portraits of how agency was distributed' is an important acknowledgment that the current data cannot characterize the distribution of voting across participants; this limitation should be moved forward into §5.2 where the voting analysis is presented, so that the reader encounters it at the point of use.","section":"§6.5"},{"comment":"Some references have incompletely formatted metadata (e.g., reference [27] begins with a URL without a title, and several URLs include raw publisher strings); these should be cleaned up for publication.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a performance-led research contribution that may fit a venue interested in design artifacts and critical HCI, though it is less conventional for a strictly empirical audience. The main technical concern—that vote-level data are interpreted as collective behavior despite unlimited voting—is fixable only by re-analyzing the data with unique-voter counts or by substantially softening the central claim. The authors' own §6.5 acknowledgment that individual-level tracking would have enabled a more detailed portrait suggests they are aware of this limit; the revision should make this limitation explicit at the point of the voting analysis, not only in the discussion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a legitimately interesting performance-led HCI paper. The artifact is new — a wearable robot on a dancer, audience voting via phones — the design process is documented in unusual detail, and the Breel-based framing of agentive behavior vs. experience of agency is applied thoughtfully. It deserves a serious referee. But the main empirical claim, that collective voting was 'strikingly consistent' and therefore subtly steered, is not actually supported by the data as reported.\n\nWhat's new: the specific system, the four-show corpus, and the honest account of how an unlimited-voting bug became a deliberate design choice. The paper situates itself well in HCI, performance studies, and interactive art, and it is transparent about the messiness of performance-led research. The survey results (150 of ~200 attendees) are plausible, and the qualitative quotes are useful.\n\nThe soft spots are real. Section 5.2 says participants could vote multiple times per prompt, and the analysis reports normalized vote ratios. A vote ratio is not a person ratio. With an average of 338–1150 votes per prompt and an audience of maybe fifty per show, a handful of hyper-voters could dominate the overrides. The paper never reports unique-voter counts. The low sigma/mu values may describe the stable behavior of a small subset, not a collective decision. The stress-test note has this right.\n\nSecond gap: there is no log of the backstage operator's actual robot commands, so the survey item 'my choices affected the robot's behavior' is never checked against what really changed. The felt-vs-exercised distinction, the heart of the paper, is missing its 'exercised' half.\n\nThat said, the authors explicitly frame the piece as a provocation, not a proof. The conceptual contribution does not depend on the consistency statistic being bulletproof — the gap between perception and verifiable influence is plausible and well-argued. My recommendation: send it to review, but ask the authors to either add unique-voter counts and a robot-command log from future shows, or soften the collective-steering language and present the voting data as illustrative. For a reading group, it is a great case study of what evidence looks like in performance-led research.","headline":"A vivid performance-led study of felt vs. exercised agency, worth reading and worth reviewing, but the voting analysis cannot distinguish collective patterns from a few hyper-voters.","tokens_in":18909,"tokens_out":2873,"would_cite":true,"duration_ms":35579,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a live audience's felt agency was decoupled from their actual power: viewers believed their votes shaped a dancer-robot duet, yet voting patterns across four performances were strikingly consistent, making the piece…","keywords":["dance","robots","human-robot interaction","interactive performances","wearables","agency","performance led research","collective agency"],"falsifier":"Conduct the same six voting prompts under a strict one-vote-per-person cap and compare the override ratios with the unlimited-voting condition; a large shift would show that repeated voting by a few individuals, not collective will, produced the consistency. Alternatively, log the robot's actual commanded motions at each prompt and check whether they match the winning option; systematic mismatches would show that the audience's choices did not change the robot's behavior at all.","tokens_in":17911,"feed_emoji":"🤖","tokens_out":7595,"duration_ms":83898,"temperature":0.7,"pith_summary":"This paper reports on Dance 2, a live performance in which audience members vote through their phones to influence a wearable robot moving on a dancer's body. Post-show surveys of 150 respondents showed that most felt they were genuinely interacting with the robot and that their choices affected the performance. Yet the logged votes across four public performances were strikingly consistent: normalized override ratios at six decision prompts varied only slightly ($\\sigma/\\mu$ between 0.034 and 0.135). The paper argues this gap shows that the audience's felt agency exceeded their actual power, and that the choreography, timing, and interface framing subtly steered collective choices. The work is offered as a live analogy for algorithmically curated digital systems where agency is felt but not exercised.","feed_headline":"Audience felt control they never actually had in live robot duet","feed_subtitle":"Across four shows, votes stayed strikingly consistent even while viewers reported real influence.","key_machinery":"The mechanism is a three-way agency loop. Audience votes are cast through a lightweight phone interface and visualized live on stage; a wearable robot based on the Calico platform travels along a silicone track stitched to the dancer's costume; and the dancer pauses, resists, or reacts to the robot's behavior, feeding cues back to the audience. The analytical device that carries the argument is the normalized override ratio at each of the six Part 4 prompts, with $\\sigma/\\mu$ used to measure cross-performance consistency. The authors deliberately kept an unlimited-votes-per-person mechanic after an early glitch, arguing that it surfaced how individuals try to amplify their voice within a collective system, while still normalizing the data to represent the group's choices.","core_discovery":"The central discovery is an empirical case of perceived agency decoupled from actual power. Audience members overwhelmingly reported that their votes influenced the robot and dancer, and two-thirds reported feeling connected to other audience members through the shared act of voting. But the consolidated voting data show stable override behavior across four performances, with the mean override ratio varying from 0.378 at prompt e to 0.846 at prompt f, and cross-performance consistency as tight as $\\sigma/\\mu = 0.034$. The paper's conclusion is that the audience's collective choices were subtly shaped by the system's design, choreography, and emotional framing, so participants experienced meaningful control even when their collective behavior followed pre-structured paths.","pith_inferences":["A direct test would re-run the same six prompts with a strict one-vote-per-person cap; if override ratios change substantially, the unlimited voting mechanic, not collective will, was driving the consistent pattern.","Logging the robot's commanded movements and comparing them against winning votes would test whether felt influence ever translated into executed behavior; the paper does not include such a log.","The same normalized-ratio method could be applied to televised audience polls and app-based town-hall votes, where participants report engagement but outcome distributions look too stable across demographics."],"forward_implications":["If the decoupling holds, interactive systems should be judged by what outcomes they actually allow users to shape, not just by whether users feel engaged.","Perceived agency can be designed into a system independently of real power, which means felt control is a designable material rather than a guarantee of influence.","The normalized override ratio offers a cheap, generalizable diagnostic for hidden steering in any shared decision-making interface, from live polls to platform recommendation votes.","The paper's framework of agentive behavior, experienced agency, and actual power gives participatory design a vocabulary for discussing the gap between those three things."],"supporting_citations":[{"why":"It supplies the distinction between agentive behavior and the experience of agency that the paper's core interpretation depends on.","marker":"[13]"},{"why":"It establishes performance-led research in the wild as the methodological frame for drawing findings from live shows.","marker":"[8]"},{"why":"It provides a baseline interactive dance work in which phone-based audience input produced varied and conflicted agency responses, an explicit contrast for the paper's design.","marker":"[4]"},{"why":"It describes the Calico wearable-robot platform whose hardware and movement system the stage robot builds on.","marker":"[35]"},{"why":"It documents how interface design shapes users' sense of agency on a major video platform, which the paper extends to its performance finding.","marker":"[28]"},{"why":"It offers a documented case of data-driven manipulation of collective behavior, invoked by the paper when connecting the result to platform-level steering.","marker":"[10]"},{"why":"It frames narrative agency in experiential theatre and supports the scaffolding design of the five-part performance structure.","marker":"[12]"}],"fun_headline_variants":["Robot duet voting: felt powerful, did little","Dance duet audience felt control they lacked","Vote illusion in live dancer-robot duet","Perceived agency, stable votes in robot duet","Audience felt influence, data says otherwise"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that normalized vote ratios represent a collective decision, even though the interface allowed unlimited votes per person, so a handful of devoted audience members could have cast many repeated votes; a second unverified assumption is that the backstage operator actually executed the winning option as specified, since no robot-behavior log accompanies the vote data.","fun_headline_variants_meta":{"raw":{"variants":["Robot duet voting: felt powerful, did little","Dance duet audience felt control they lacked","Vote illusion in live dancer-robot duet","Perceived agency, stable votes in robot duet","Audience felt influence, data says otherwise"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000142,"raw_usage":{"total_tokens":1112,"prompt_tokens":830,"completion_tokens":282,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":446,"completion_tokens_details":{"reasoning_tokens":209}},"tokens_in":446,"tokens_out":282,"duration_ms":4203,"temperature":1.0,"reasoning_tokens":209,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:35:18.042018+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct the same six voting prompts under a strict one-vote-per-person cap and compare the override ratios with the unlimited-voting condition; a large shift would show that repeated voting by a few individuals, not collective will, produced the consistency. Alternatively, log the robot's actual commanded motions at each prompt and check whether they match the winning option; systematic mismatches would show that the audience's choices did not change the robot's behavior at all.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the distinction between agentive behavior and the experience of agency that the paper's core interpretation depends on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It establishes performance-led research in the wild as the methodological frame for drawing findings from live shows."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It describes the Calico wearable-robot platform whose hardware and movement system the stage robot builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It offers a documented case of data-driven manipulation of collective behavior, invoked by the paper when connecting the result to platform-level steering."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It frames narrative agency in experiential theatre and supports the scaffolding design of the five-part performance structure."}],"review_version":1}