REVIEW 3 major objections 4 minor 11 references
Clones in the Machine: A Feminist Critique of Agency in Digital Cloning
T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Researcher-driven digital cloning should be governed like research with human participants.
desk verdict A short, well-written position paper that makes a plausible ethical case for consent in digital cloning research, but its policy conclusion rests on an unargued premise about what clones actually simulate. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the 'digital clone'—a computational agent that replicates a user's digital history, such as posts and interactions, to simulate and predict behavior. The load-bearing mechanism is a feminist relational account of agency, which holds that agency does not reside in discrete agents but emerges from sociomaterial arrangements; on this view, a clone built from a user's data functions as an extension of that user's agency. This substitution—data replica as agentic extension—is what converts a privacy concern into a consent-and-representation concern that warrants human-subjects oversight.
What would settle it
Conduct a controlled comparison in which a digital clone and a demographic baseline each predict a user's responses in context-sensitive situations, such as expressing political views in private versus public settings. If the clone's predictions are no more accurate or context-sensitive than the baseline, then clones have not been shown to simulate agency rather than aggregate patterns, and the paper's central premise is empirically weakened.
Extended reading notes
Core claim
The central claim is that digital clones, as used in simulation studies, simulate user agency and therefore cloning is not a form of passive data collection. Because agency is relational and context-dependent—shaped by power structures, gender, and social position—the paper argues that clones which strip away context misrepresent users, especially marginalized users, and can perpetuate systemic biases. Researcher-driven cloning that scrapes public data without explicit consent therefore fails the ethical standards expected of research with human participants. The paper concludes that digital cloning research should be subject to human-participant ethical approval, and recommends governance mechanisms such as dynamic consent and decentralized, non-commercial data repositories.
Load-bearing premise
The entire argument depends on the premise that digital clones do more than aggregate data—that they actively simulate behavior and agency; if clones are merely statistical pattern-matchers with no meaningful simulation of agency, the distinct harm the paper identifies collapses into a generic privacy concern.
Editorial extensions
If this is right
- Research ethics boards would classify researcher-driven digital cloning as human-subjects research, requiring informed consent even when the source data is publicly available.
- Simulation studies that clone users would have to justify any waiver of explicit consent, or switch to user-driven donation repositories with transparent terms.
- Persistent clones would need deletion and withdrawal mechanisms, honoring users' right to be forgotten after they revoke consent.
- Dynamic consent dashboards—notifying users of data use, outcomes, and withdrawal options—would become standard infrastructure for behavior-simulation research.
Reading between the lines
- The relational-agency premise, if accepted, extends beyond academic simulation studies to commercial LLM-persona research, implying that companies building user simulations from public text would also face a consent obligation.
- A testable extension of the proposed governance is whether participants in voluntary donation repositories behave differently from users whose data is scraped, which would indicate whether consent changes the fidelity of simulated behavior.
- The argument implies a legal corollary: a clone is a new derived artifact, not the original data, so existing data-protection regimes may already require consent for cloning even where scraping the raw data is lawful.
- If the paper is right, the ethical burden also shifts to the designers of simulation frameworks, who would need to build context-awareness and consent mechanisms into the tools themselves, not just the studies that use them.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a four-page position paper arguing that researcher-driven digital cloning—the creation of computational agents that replicate users' digital histories to simulate and predict behavior—raises ethical concerns that are not captured by existing data-protection and human-subjects frameworks. Drawing on feminist theory (Suchman, Toupin, Nunes), the author contends that clones are not merely aggregates of data but active simulations of user agency, and that using public data to build them without explicit consent misrepresents users, flattens contextual complexity, and risks reinforcing systemic biases. The paper proposes two remedies: decentralized, ethically governed data-donation repositories and dynamic consent models with participatory dashboards. It concludes that digital cloning research should be subject to human-participant ethical approval.
Significance. If accepted, the paper's central recommendation would require HCI and adjacent fields to classify simulation studies that clone users as human-subjects research, with concrete consequences for IRB review, data governance, and consent infrastructure. The paper is clearly written and well-sourced, and it connects a timely methodological practice (LLM-driven personas and agent-based simulations) to a substantive ethical debate. Its main strength is its synthesis of feminist critiques of agency into a concrete call for procedural change. However, the argument is built on a factual premise about what digital clones do that is not specified or empirically supported, and the proposed policy intervention lacks a clear scope. The paper is an honest position piece, but the central claim needs to be sharpened before the recommendation can be assessed.
major comments (3)
- [Sections 1 and 2] The claim that digital clones 'do more than aggregate data—they simulate behaviors and agency' is load-bearing for the entire argument, but the paper gives no operational definition of agency and no direct evidence that the cited systems (Puri et al. [6], Schmidt et al. [7]) instantiate it. If clones are only sophisticated statistical predictors, the specific harms described (identity misrepresentation, manipulation, persistent persona) reduce to familiar privacy or bias concerns, and the proposed human-participant review becomes either unmotivated or overbroad. Please specify the minimal factual premise required for the ethical argument—e.g., that clones are interactive, persistent, and identifiable—and either defend that premise in the cited examples or reframe the argument so it does not depend on the contested term 'agency.'
- [Sections 2 and 6] The central recommendation—'Digital cloning research should be exposed to human-participant ethical approval'—presupposes a clear boundary between digital cloning and ordinary analysis of public data. The paper does not define what counts as a digital clone for regulatory purposes, beyond saying it 'actively simulate[s] user agency.' Without an operational boundary, the recommendation either covers all machine learning on public user data or is unenforceable. Please provide criteria (e.g., interactivity, persistence, identifiability, use in simulation) that distinguish cloning from standard data mining and make the scope of the policy concrete.
- [Section 5] The proposal of dynamic consent models for digital cloning does not address the feasibility of retrospective consent for existing large-scale datasets. Puri et al. [6], the paper's central example, cloned histories of over 10,000 users scraped from public platforms; it is not explained how researchers would contact all users to obtain consent or what should happen to already-built clones if consent is withheld. Without a transition plan for existing data, the proposal is incomplete as a policy recommendation. Please address how dynamic consent would apply retrospectively.
minor comments (4)
- [Title and author block] The title contains 'Digit al Cloning' and the affiliation 'Univers ity of Amsterdam'; these spacing artifacts should be fixed in the camera-ready version.
- [Section 2] The description of Puri et al. [6] states that the ethics board deemed explicit consent unnecessary due to public availability, but the paper does not say whether the user data was anonymized before cloning; this detail matters because the argument assumes clones are linked to identifiable individuals.
- [Section 4] The claim that clones 'may overemphasize frequently repeated behaviors while overlooking passive or evolving user interactions, reinforcing echo chambers and distorting online discourse' is an empirical hypothesis without citation; it should be explicitly marked as a projection or supporting example rather than a demonstrated effect.
- [Section 5] The 'decentralized data donation repositories' are described as operating 'exclusively for non-commercial academic research,' but the paper does not discuss governance, funding, or how community oversight would be enforced; a sentence on these practicalities would strengthen the proposal.
Circularity Check
No significant circularity: this is a position paper whose normative argument rests on external, cited feminist theory and a stipulative definition, not on fitted predictions or self-referential derivation.
full rationale
Score 0. This paper is an explicitly self-described position paper ('This position paper examines these ethical challenges', Section 1), not a derivation of empirical predictions from fitted parameters. Its load-bearing premise—that digital clones 'simulate behaviors and agency'—is imported from external sources ([6] Puri et al., [7] Schmidt et al.) and from the paper's own stipulative definition of digital clones as 'computational agents that replicate users' digital histories (e.g., posts, interactions) to simulate and predict behaviors'. A stipulative definition or an external citation may be contestable, but it is not circular: the paper never fits a parameter and then presents that parameter as a prediction, and the reference list contains no self-citations by the author. The policy conclusion in Section 6 follows by normative argument from feminist theory ([8], [9], [10]), which is genuinely external to the paper. Whether clones actually satisfy the definition's agency-simulation predicate is a factual-support weakness—correctly flagged by the skeptic as an unsupported premise—but the specific reduction-to-inputs pattern required for circularity is absent. No equation is fitted from data and then renamed as a result; no uniqueness claim is imported from the author's prior work; and no known empirical pattern is merely relabeled.
Assumptions & free parameters
assumptions (4)
- domain assumption Agency is relational and shaped by power structures.
- domain assumption Digital footprints are extensions of the self.
- domain assumption Publicly available data used in simulation constitutes a morally significant extension of the person.
- domain assumption Digital clones simulate behaviors and agency, not merely aggregate data.
Cite this review
Pith. "Pith review of Clones in the Machine: A Feminist Critique of Agency in Digital Cloning." pith.science (2026). https://pith.science/paper/MW23IWX2
@misc{pith2026250418807,
author = {Pith},
title = {Pith review of: Clones in the Machine: A Feminist Critique of Agency in Digital Cloning},
year = {2026},
howpublished = {\url{https://pith.science/paper/MW23IWX2}},
note = {Machine review of arXiv:2504.18807}
}
read the original abstract
This paper critiques digital cloning in academic research, highlighting how it exemplifies AI solutionism. Digital clones, which replicate user data to simulate behavior, are often seen as scalable tools for behavioral insights. However, this framing obscures ethical concerns around consent, agency, and representation. Drawing on feminist theories of agency, the paper argues that digital cloning oversimplifies human complexity and risks perpetuating systemic biases. To address these issues, it proposes decentralized data repositories and dynamic consent models, promoting ethical, context-aware AI practices that challenge the reductionist logic of AI solutionism
Reference graph
Works this paper leans on
-
[6]
Prateek Puri, Gabriel Hassler, Sai Katragadda, and Anto n Shenk. 2024. Digital cloning of online social networks for language-sensitive agent-based modeling of misinformation spread. PLOS ONE 19, 6 (June 2024), e0304889. doi:10.1371/journal.pone.0304889 Publisher: Public Library of Science
-
[7]
Albrecht Schmidt, Passant Elagroudy, Fiona Draxler, Fr auke Kreuter, and Robin Welsch. 2024. Simulating the Human i n HCD with ChatGPT: Redesigning Interaction Design with AI. Interactions 31, 1 (Jan. 2024), 24–31. doi:10.1145/3637436
doi:10.1145/3637436 2024
-
[1]
Yida Chen, Aoyu Wu, Trevor DePodesta, Catherine Yeh, Ken neth Li, Nicholas Castillo Marin, Oam Patel, Jan Riecke, Shi vam Raval, Olivia Seow, Martin Wattenberg, and Fernanda Viégas. 2024. Design ing a Dashboard for Transparency and Control of Conversatio nal AI. doi:10.48550/arXiv.2406.07882 arXiv:2406.07882 [cs]
- [2]
-
[3]
Casey Fiesler. 2020. Lawful Users: Copyright Circumven tion and Legal Constraints on Technology Use. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . ACM, Honolulu HI USA, 1–11. doi:10.1145/3313831.3376745
arXiv 2020
-
[4]
Stamatis Karnouskos. 2020. Artificial Intelligence in D igital Media: The Era of Deepfakes. IEEE Transactions on Technology and Society 1, 3 (Sept. 2020), 138–147. doi:10.1109/TTS.2020.3001312 Conference Name: IEEE Transactions on Technology and Socie ty
arXiv 2020
-
[5]
Rafaela Nunes. 2024. AI, a Tool or an Author? A Posthuman F eminist Perspective on the Agency of Gen-AI in Creative Prac tices. Augmented Human Research 9, 1 (Dec. 2024), 8. doi:10.1007/s41133-024-00074-8
-
[8]
Becca Schwartz and Gina Neff. 2019. The gendered affordanc es of Craigslist “new-in-town girls wanted” ads. New Media & Society 21, 11-12 (Nov. 2019), 2404–2421. doi:10.1177/1461444819849897
Show all 11 references
-
[9]
Lucy Suchman. 2020. Agencies in Technology Design: Femi nist Reconfigurations*. In Machine Ethics and Robot Ethics (1 ed.), Wendell Wallach and Peter Asaro (Eds.). Routledge, 361–375. doi:10.4324/9781003074991-32
2020 doi
-
[10]
Sophie Toupin. 2024. Shaping feminist artificial intel ligence. New Media & Society 26, 1 (Jan. 2024), 580–595. doi:10.1177/14614448221150776 Publisher: SAGE Publications
2024 doi
-
[11]
acm-jdslogo.png
Jon Truby and Rafael Brown. 2021. Human digital thought clones: the Holy Grail of artificial intelligence for big data. Information & Communications Technology Law 30, 2 (May 2021), 140–168. doi:10.1080/13600834.2020.1850174 Publisher: Routledge. Manuscript submitted to ACM Thi...
2021
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.