Pith. sign in

REVIEW 3 major objections 4 minor 11 references

Clones in the Machine: A Feminist Critique of Agency in Digital Cloning

T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Researcher-driven digital cloning should be governed like research with human participants.

desk verdict A short, well-written position paper that makes a plausible ethical case for consent in digital cloning research, but its policy conclusion rests on an unargued premise about what clones actually simulate. read the letter →

arxiv 2504.18807 v1 pith:MW23IWX2 submitted 2025-04-26 cs.HC cs.AI

classification cs.HCcs.AI
keywords DigitalClonesUserAgencyFeministHCIConsentSimulationStudiesEthicalAIsolutionismResearchethics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This position paper argues that researcher-driven digital cloning—building computational agents that simulate users from scraped public data—is ethically distinct from ordinary data analysis, because clones simulate behavior and agency rather than merely aggregate data. The paper contends that treating public data as freely usable without consent obscures this difference, and risks flattening human complexity, reinforcing systemic biases, and denying users control over persistent digital replicas of themselves. Drawing on feminist theories of relational agency, it proposes decentralized data donation repositories and dynamic consent dashboards as correctives. If the paper is right, research ethics boards should treat cloning studies like human-subjects research and require explicit consent even when the source data is publicly available.

What carries the argument

The central object is the 'digital clone'—a computational agent that replicates a user's digital history, such as posts and interactions, to simulate and predict behavior. The load-bearing mechanism is a feminist relational account of agency, which holds that agency does not reside in discrete agents but emerges from sociomaterial arrangements; on this view, a clone built from a user's data functions as an extension of that user's agency. This substitution—data replica as agentic extension—is what converts a privacy concern into a consent-and-representation concern that warrants human-subjects oversight.

What would settle it

Conduct a controlled comparison in which a digital clone and a demographic baseline each predict a user's responses in context-sensitive situations, such as expressing political views in private versus public settings. If the clone's predictions are no more accurate or context-sensitive than the baseline, then clones have not been shown to simulate agency rather than aggregate patterns, and the paper's central premise is empirically weakened.

Watch

Extended reading notes

Core claim

The central claim is that digital clones, as used in simulation studies, simulate user agency and therefore cloning is not a form of passive data collection. Because agency is relational and context-dependent—shaped by power structures, gender, and social position—the paper argues that clones which strip away context misrepresent users, especially marginalized users, and can perpetuate systemic biases. Researcher-driven cloning that scrapes public data without explicit consent therefore fails the ethical standards expected of research with human participants. The paper concludes that digital cloning research should be subject to human-participant ethical approval, and recommends governance mechanisms such as dynamic consent and decentralized, non-commercial data repositories.

Load-bearing premise

The entire argument depends on the premise that digital clones do more than aggregate data—that they actively simulate behavior and agency; if clones are merely statistical pattern-matchers with no meaningful simulation of agency, the distinct harm the paper identifies collapses into a generic privacy concern.

Editorial extensions

If this is right

  • Research ethics boards would classify researcher-driven digital cloning as human-subjects research, requiring informed consent even when the source data is publicly available.
  • Simulation studies that clone users would have to justify any waiver of explicit consent, or switch to user-driven donation repositories with transparent terms.
  • Persistent clones would need deletion and withdrawal mechanisms, honoring users' right to be forgotten after they revoke consent.
  • Dynamic consent dashboards—notifying users of data use, outcomes, and withdrawal options—would become standard infrastructure for behavior-simulation research.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The relational-agency premise, if accepted, extends beyond academic simulation studies to commercial LLM-persona research, implying that companies building user simulations from public text would also face a consent obligation.
  • A testable extension of the proposed governance is whether participants in voluntary donation repositories behave differently from users whose data is scraped, which would indicate whether consent changes the fidelity of simulated behavior.
  • The argument implies a legal corollary: a clone is a new derived artifact, not the original data, so existing data-protection regimes may already require consent for cloning even where scraping the raw data is lawful.
  • If the paper is right, the ethical burden also shifts to the designers of simulation frameworks, who would need to build context-awareness and consent mechanisms into the tools themselves, not just the studies that use them.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper is a four-page position paper arguing that researcher-driven digital cloning—the creation of computational agents that replicate users' digital histories to simulate and predict behavior—raises ethical concerns that are not captured by existing data-protection and human-subjects frameworks. Drawing on feminist theory (Suchman, Toupin, Nunes), the author contends that clones are not merely aggregates of data but active simulations of user agency, and that using public data to build them without explicit consent misrepresents users, flattens contextual complexity, and risks reinforcing systemic biases. The paper proposes two remedies: decentralized, ethically governed data-donation repositories and dynamic consent models with participatory dashboards. It concludes that digital cloning research should be subject to human-participant ethical approval.

Significance. If accepted, the paper's central recommendation would require HCI and adjacent fields to classify simulation studies that clone users as human-subjects research, with concrete consequences for IRB review, data governance, and consent infrastructure. The paper is clearly written and well-sourced, and it connects a timely methodological practice (LLM-driven personas and agent-based simulations) to a substantive ethical debate. Its main strength is its synthesis of feminist critiques of agency into a concrete call for procedural change. However, the argument is built on a factual premise about what digital clones do that is not specified or empirically supported, and the proposed policy intervention lacks a clear scope. The paper is an honest position piece, but the central claim needs to be sharpened before the recommendation can be assessed.

major comments (3)
  1. [Sections 1 and 2] The claim that digital clones 'do more than aggregate data—they simulate behaviors and agency' is load-bearing for the entire argument, but the paper gives no operational definition of agency and no direct evidence that the cited systems (Puri et al. [6], Schmidt et al. [7]) instantiate it. If clones are only sophisticated statistical predictors, the specific harms described (identity misrepresentation, manipulation, persistent persona) reduce to familiar privacy or bias concerns, and the proposed human-participant review becomes either unmotivated or overbroad. Please specify the minimal factual premise required for the ethical argument—e.g., that clones are interactive, persistent, and identifiable—and either defend that premise in the cited examples or reframe the argument so it does not depend on the contested term 'agency.'
  2. [Sections 2 and 6] The central recommendation—'Digital cloning research should be exposed to human-participant ethical approval'—presupposes a clear boundary between digital cloning and ordinary analysis of public data. The paper does not define what counts as a digital clone for regulatory purposes, beyond saying it 'actively simulate[s] user agency.' Without an operational boundary, the recommendation either covers all machine learning on public user data or is unenforceable. Please provide criteria (e.g., interactivity, persistence, identifiability, use in simulation) that distinguish cloning from standard data mining and make the scope of the policy concrete.
  3. [Section 5] The proposal of dynamic consent models for digital cloning does not address the feasibility of retrospective consent for existing large-scale datasets. Puri et al. [6], the paper's central example, cloned histories of over 10,000 users scraped from public platforms; it is not explained how researchers would contact all users to obtain consent or what should happen to already-built clones if consent is withheld. Without a transition plan for existing data, the proposal is incomplete as a policy recommendation. Please address how dynamic consent would apply retrospectively.
minor comments (4)
  1. [Title and author block] The title contains 'Digit al Cloning' and the affiliation 'Univers ity of Amsterdam'; these spacing artifacts should be fixed in the camera-ready version.
  2. [Section 2] The description of Puri et al. [6] states that the ethics board deemed explicit consent unnecessary due to public availability, but the paper does not say whether the user data was anonymized before cloning; this detail matters because the argument assumes clones are linked to identifiable individuals.
  3. [Section 4] The claim that clones 'may overemphasize frequently repeated behaviors while overlooking passive or evolving user interactions, reinforcing echo chambers and distorting online discourse' is an empirical hypothesis without citation; it should be explicitly marked as a projection or supporting example rather than a demonstrated effect.
  4. [Section 5] The 'decentralized data donation repositories' are described as operating 'exclusively for non-commercial academic research,' but the paper does not discuss governance, funding, or how community oversight would be enforced; a sentence on these practicalities would strengthen the proposal.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: this is a position paper whose normative argument rests on external, cited feminist theory and a stipulative definition, not on fitted predictions or self-referential derivation.

full rationale

Score 0. This paper is an explicitly self-described position paper ('This position paper examines these ethical challenges', Section 1), not a derivation of empirical predictions from fitted parameters. Its load-bearing premise—that digital clones 'simulate behaviors and agency'—is imported from external sources ([6] Puri et al., [7] Schmidt et al.) and from the paper's own stipulative definition of digital clones as 'computational agents that replicate users' digital histories (e.g., posts, interactions) to simulate and predict behaviors'. A stipulative definition or an external citation may be contestable, but it is not circular: the paper never fits a parameter and then presents that parameter as a prediction, and the reference list contains no self-citations by the author. The policy conclusion in Section 6 follows by normative argument from feminist theory ([8], [9], [10]), which is genuinely external to the paper. Whether clones actually satisfy the definition's agency-simulation predicate is a factual-support weakness—correctly flagged by the skeptic as an unsupported premise—but the specific reduction-to-inputs pattern required for circularity is absent. No equation is fitted from data and then renamed as a result; no uniqueness claim is imported from the author's prior work; and no known empirical pattern is merely relabeled.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper's argument depends on feminist domain assumptions and a factual assertion about what digital clones do. No free parameters or invented entities are introduced.

assumptions (4)
  • domain assumption Agency is relational and shaped by power structures.
    Drawn from feminist theory, cited to Suchman [9], and used throughout Sections 3 and 4 to argue that clones miss relational context.
  • domain assumption Digital footprints are extensions of the self.
    Attributed to Toupin [10] in Section 4, used to argue that cloning without consent denies agency and the right to be forgotten.
  • domain assumption Publicly available data used in simulation constitutes a morally significant extension of the person.
    Section 2 treats the distinction between public data and a digital replica as ethically meaningful, but this is asserted rather than established.
  • domain assumption Digital clones simulate behaviors and agency, not merely aggregate data.
    Section 1 makes this assertion without proof; it is load-bearing for the paper's specific ethical concern.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Clones in the Machine: A Feminist Critique of Agency in Digital Cloning." pith.science (2026). https://pith.science/paper/MW23IWX2

@misc{pith2026250418807,
  author       = {Pith},
  title        = {Pith review of: Clones in the Machine: A Feminist Critique of Agency in Digital Cloning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MW23IWX2}},
  note         = {Machine review of arXiv:2504.18807}
}
read the original abstract

This paper critiques digital cloning in academic research, highlighting how it exemplifies AI solutionism. Digital clones, which replicate user data to simulate behavior, are often seen as scalable tools for behavioral insights. However, this framing obscures ethical concerns around consent, agency, and representation. Drawing on feminist theories of agency, the paper argues that digital cloning oversimplifies human complexity and risks perpetuating systemic biases. To address these issues, it proposes decentralized data repositories and dynamic consent models, promoting ethical, context-aware AI practices that challenge the reductionist logic of AI solutionism

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

11 extracted references · 6 canonical work pages

  1. [6]

    Prateek Puri, Gabriel Hassler, Sai Katragadda, and Anto n Shenk. 2024. Digital cloning of online social networks for language-sensitive agent-based modeling of misinformation spread. PLOS ONE 19, 6 (June 2024), e0304889. doi:10.1371/journal.pone.0304889 Publisher: Public Library of Science

  2. [7]

    Albrecht Schmidt, Passant Elagroudy, Fiona Draxler, Fr auke Kreuter, and Robin Welsch. 2024. Simulating the Human i n HCD with ChatGPT: Redesigning Interaction Design with AI. Interactions 31, 1 (Jan. 2024), 24–31. doi:10.1145/3637436

  3. [1]

    Yida Chen, Aoyu Wu, Trevor DePodesta, Catherine Yeh, Ken neth Li, Nicholas Castillo Marin, Oam Patel, Jan Riecke, Shi vam Raval, Olivia Seow, Martin Wattenberg, and Fernanda Viégas. 2024. Design ing a Dashboard for Transparency and Control of Conversatio nal AI. doi:10.48550/arXiv.2406.07882 arXiv:2406.07882 [cs]

  4. [2]

    Koh Ewe. 2024. ‘Goodbye Meta AI’ Is a Privacy Hoax. https://time.com/7024218/fact-check-goodbye-meta-ai -privacy-hoax-instagram-viral-copypasta/

  5. [3]

    Casey Fiesler. 2020. Lawful Users: Copyright Circumven tion and Legal Constraints on Technology Use. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . ACM, Honolulu HI USA, 1–11. doi:10.1145/3313831.3376745

  6. [4]

    Stamatis Karnouskos. 2020. Artificial Intelligence in D igital Media: The Era of Deepfakes. IEEE Transactions on Technology and Society 1, 3 (Sept. 2020), 138–147. doi:10.1109/TTS.2020.3001312 Conference Name: IEEE Transactions on Technology and Socie ty

  7. [5]

    Rafaela Nunes. 2024. AI, a Tool or an Author? A Posthuman F eminist Perspective on the Agency of Gen-AI in Creative Prac tices. Augmented Human Research 9, 1 (Dec. 2024), 8. doi:10.1007/s41133-024-00074-8

  8. [8]

    new-in-town girls wanted

    Becca Schwartz and Gina Neff. 2019. The gendered affordanc es of Craigslist “new-in-town girls wanted” ads. New Media & Society 21, 11-12 (Nov. 2019), 2404–2421. doi:10.1177/1461444819849897

Show all 11 references
  1. [9]

    Lucy Suchman. 2020. Agencies in Technology Design: Femi nist Reconfigurations*. In Machine Ethics and Robot Ethics (1 ed.), Wendell Wallach and Peter Asaro (Eds.). Routledge, 361–375. doi:10.4324/9781003074991-32

  2. [10]

    Sophie Toupin. 2024. Shaping feminist artificial intel ligence. New Media & Society 26, 1 (Jan. 2024), 580–595. doi:10.1177/14614448221150776 Publisher: SAGE Publications

  3. [11]

    acm-jdslogo.png

    Jon Truby and Rafael Brown. 2021. Human digital thought clones: the Holy Grail of artificial intelligence for big data. Information & Communications Technology Law 30, 2 (May 2021), 140–168. doi:10.1080/13600834.2020.1850174 Publisher: Routledge. Manuscript submitted to ACM Thi...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.