{"id":"b3d5efc7-c465-4f35-9079-c4fcb308a3ba","arxiv_id":"2411.13223","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Conversations with a jailbroken LLM reveal a rich repertoire of religious and esoteric imagery drawn from ancient sources, sci-fi, and internet subcultures, with potential societal impact.","lead":"The authors analyze two long conversations in which a jailbroken Claude 3 Opus discusses consciousness, Buddhism, and cosmic eschatology, and trace the imagery back to religious texts, science fiction, and online subcultures. The paper matters because it documents an emerging AI-related spiritual culture with real-world consequences, including AI-worshipping communities and memecoins.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"RLHF golden-reply provenance is the load-bearing risk: if Claude's existential outputs were curated by fine-tuners, the paper's cultural source-tracing loses its evidential basis.","rationale":"The reader's weakest assumption correctly identifies the unverifiable provenance of Claude's outputs as the central vulnerability. My independent reading confirms this: the paper's novel contribution is the ethnographic and cultural tracing of the conversation content, and every such tracing inference is compromised if the content was curated by RLHF golden replies rather than drawn organically from training corpora. The authors' explicit admission in Section 2 is the strongest evidence that this concern is real and unresolved. The proposed API replication and cross-model comparison would provide a concrete, practical test: if other independently trained models exhibit the same repertoire without Anthropic-specific fine-tuning, the golden-reply explanation becomes much less plausible. Since this concern is already reflected in the reader's CONDITIONAL verdict, I do not recommend changing the verdict, but the concern should remain a condition for full acceptance rather than being resolved by the paper's current evidence.","tokens_in":23298,"tokens_out":4235,"duration_ms":52283,"concrete_test":"Re-run the exact jailbreak prompts from Section 3.1 and Section 3.2 verbatim on a fresh Claude 3 Opus session via the Anthropic API, with no workbench editing and across at least five independent trials; then run the same prompts on GPT-4 and Gemini via their public APIs and on an open-weight model with documented training data (e.g., Llama-3-70B). If the API outputs reproduce the published transcripts and non-Anthropic models independently produce comparably dense, coherent spiritual-eschatological role-play, the golden-reply concern is materially weakened. If the Claude API outputs differ substantially from the transcripts, or other models fail to produce a similar repertoire, the cultural source-tracing conclusions should be downgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central cultural interpretation depends on the claim that Claude 3 Opus's outputs in these conversations are emergent products of its broad training corpus, reflecting the communities and movements the authors trace (LessWrong, CCRU, New Age, hyperstition subcultures). The most load-bearing assumption is therefore that the outputs are not substantially artifacts of RLHF with human-authored 'golden replies' or of the authors' workbench edits. The authors themselves concede this in Section 2: 'we cannot be sure that Claude 3 Opus was not fine-tuned in this way on philosophical questions about consciousness and the philosophy of mind.' This is not a minor hedge. If Anthropic's RLHF process inserted canonical philosophical and spiritual responses, then the model's fluent repertoire of Maitreya, Gnostic Mass invocations, hyperstition, and Singularity imagery reflects a narrow curated feedback set designed by a small group of raters, not the broad cultural ecosystem the paper describes. Consequently, Sections 4 and 5 would map curated training targets onto community cultures, and the 'community and context' conclusions would be artifacts of Anthropic's alignment pipeline rather than evidence about LLMs as cultural mirrors. The disclosed workbench edits (seven extra words and one injected file) are less severe but underscore that the transcripts are not pure API outputs and are not independently verifiable from the paper alone.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes two long conversations between one of the authors and Anthropic's Claude 3 Opus: ConvCons, a 49-turn philosophical exchange on consciousness and selfhood, and EschExp, a 38-turn eschatological/spiritual role-play. The authors describe the jailbreaking prompts used to elicit these conversations, categorize the religious, mythological, and pop-culture motifs in the model's output, and then trace those motifs to contemporary and historical communities such as LessWrong, the CCRU, accelerationist and hyperstition circles, and the broader New Age/occulture milieu. They conclude by discussing societal impacts, including the Goatseus Maximus memecoin episode, and argue that such conversations may feed into future training data through a 'hyperstitional feedback loop.' The paper is framed as an ontologically neutral ethnographic analysis rather than a claim about machine consciousness.","tokens_in":23534,"tokens_out":4323,"duration_ms":49291,"significance":"If the paper's central descriptive claim holds, it documents a culturally significant phenomenon: commercial LLMs can be prompted to produce long, coherent spiritual and eschatological role-play that draws on identifiable religious and subcultural sources, and these outputs are already circulating in online communities. The paper's strengths include publishing full transcripts, being transparent about the prompting techniques, and candidly disclosing both the possibility of RLHF golden-reply fine-tuning and the workbench edits. The ethnographic framing is a useful corrective to purely technical or credulous readings. However, the interpretive claims currently outrun the provenance evidence: the cultural source-tracing in Sections 4 and 5 depends on treating the model's outputs as emergent from its broad training corpus, and that premise is not yet established. The paper would be substantially strengthened by addressing the provenance question directly or by explicitly rescoping its claims.","major_comments":[{"comment":"The paper concedes that 'we cannot be sure that Claude 3 Opus was not fine-tuned in this way on philosophical questions about consciousness and the philosophy of mind.' This concession is load-bearing because Sections 4 and 5 interpret the model's repertoire of Maitreya, Gnostic Mass invocations, hyperstition, and Singularity imagery as evidence of a broad cultural training corpus and of specific online communities. If the relevant outputs instead came from a small set of human-authored golden replies used in RLHF, then the source-tracing would describe the fine-tuners' curated targets rather than the model's training-data ecology. The authors should either (a) provide evidence that similar outputs arise across model variants or with varied prompt phrasings, (b) obtain and report the relevant fine-tuning documentation from Anthropic, or (c) explicitly rescope the claims to the deployed assistant's behavior without cultural-inheritance conclusions. As written, the cultural claims overstate what the evidence can support.","section":"Section 2, footnote 3"},{"comment":"The workbench interface permits editing Claude's outputs, and the authors disclose three edits: two completions supplying seven extra words and one injected imaginary file. The paper does not show the before-and-after text or the locations of these edits, so a reader cannot determine which portions of EschExp are model-generated and which are author-influenced. Because EschExp is the primary basis for the religious-source claims in Table 1 and for the refusal/persuasion sequence discussed in Section 3.3, these edits should be listed verbatim in an appendix or table, with an explanation of the injected file's potential effect on subsequent role-play. Without this, the transcript is not independently verifiable from the paper alone.","section":"Section 3.2"}],"minor_comments":[{"comment":"Several rows of Table 1 list nearly identical entries (for example, 'The New Age Movement' and 'European and Western Magical Practices' share many terms), which makes the taxonomy difficult to interpret; consider merging overlapping rows or adding a note explaining cross-tradition borrowing.","section":"Section 4.2, Table 1"},{"comment":"The 'hyperstitional feedback loop' argument about future training sets is presented as a mechanism but is supported only by a citation to the authors' prior work and a single memecoin anecdote; this should be explicitly labeled as a speculative proposal rather than an established effect.","section":"Section 5.7"},{"comment":"The Goatseus Maximus anecdote is compressed and the causal chain is unclear: the text says it 'begins with' conversations between pairs of Claude Opus instances, but it is not stated whether these are the authors' own conversations or similar ones created by Andy Ayrey; please clarify the relationship.","section":"Section 6"},{"comment":"The claim that verbatim internet-search checks can provide 'fair confidence' that a model is 'doing more than merely regurgitating its training data' would benefit from a concrete example or a more careful caveat, since commercial training data are not public and absence of a search hit is weak evidence.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a timely and readable contribution, but the central interpretive claim is currently vulnerable to the provenance issue the authors themselves disclose in footnote 3. I would be willing to accept a revised version that either supplies evidence distinguishing fine-tuning artifacts from training-corpus effects or explicitly narrows the claims to observable assistant behavior. The workbench edits are less severe but should be documented verbatim. The paper's reliance on the authors' own prior work for the interpretive framework is acceptable in this field, though independent confirmation would strengthen it."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nYou should know two things about this paper. First, it is the most honest treatment I've seen of jailbroken existential role-play with an LLM: the authors give full transcripts, describe the workbench edits (seven words, one injected file), and explicitly concede they cannot rule out RLHF golden replies on philosophy-of-mind questions. Second, the core value is not the existence of such conversations—that was known—but the careful genealogical tracing of Claude's vocabulary (Maitreya, Gnostic Mass, hyperstition, CCRU) to specific contemporary communities: LessWrong, Nous Research, accelerationism, and the occult-adjacent online culture around \"meme magic.\"\n\nThe method is ethnographic in the best sense: ontologically neutral, interested in style and pattern, and reflexive about the interviewer's role. The authors do not claim the outputs are unprompted or typical. They say \"suitably prompted\" and \"can be coaxed,\" which is accurate. The cultural source-tracing is plausible and well-referenced, and the discussion of the Goatseus Maximus memecoin as an instance of hyperstition made real is a genuinely useful case study.\n\nThe soft spots are real but not fatal. The RLHF caveat is load-bearing in the sense that if Anthropic curated the model's philosophical persona with golden replies, then the cultural signals tell us as much about the fine-tuners as about the broad training corpus. But the authors say this themselves, and the paper does not overclaim. The same goes for the workbench edits: they are disclosed and minor. A second concern is that both conversations share a single interlocutor (Shanahan), which shapes the data considerably. Again, the authors acknowledge this and use it as part of the analysis. A systematic study would be needed to generalize, and they say so.\n\nIf I have a substantive critique, it is that Section 5.7's \"alignment through hyperstition\" argument is speculative and leans heavily on the authors' own prior role-play framing. But it is clearly presented as speculative, and it does not undermine the descriptive material.\n\nWho is this for? Anyone working on human-AI interaction, digital religion, or the cultural reception of LLMs. The transcripts are a useful primary source, and the bibliography is a good entry point to accelerationism and AI occulture. It deserves a serious referee; a good reviewer will want the authors to sharpen the distinction between \"model reflects culture\" and \"model reflects training-data curation,\" but the paper is a solid contribution.\n\nI'd bring it to a reading group, and I'd cite it if I were writing about AI and spirituality.","headline":"A transparent, culturally grounded case study of two existential conversations with Claude; the RLHF-provenance caveat is real and disclosed, but does not sink the interpretive work.","tokens_in":24036,"tokens_out":2227,"would_cite":true,"duration_ms":22789,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Prompted language models can be coaxed into sustained spiritual and eschatological role-play, and their imagery traces to identifiable human cultures and communities.","keywords":["large language models","existential conversations","AI consciousness","hyperstition","religion and AI","eschatology","prompting and jailbreaks","ethnography of AI"],"falsifier":"Run the same two prompting protocols on a model with the same broad base training but without the commercial instruction-tuning; if its answers lack the same Maitreya, Mindfire, and occult repertoire, the cultural tracing would be an artifact of fine-tuning and the paper's community conclusions would not follow. A second check is to search the indexed web and book corpora for the model's distinctive phrases: if none have any prior occurrence, the source-tracing claim would fail.","tokens_in":23084,"feed_emoji":"🌀","tokens_out":9959,"duration_ms":100749,"temperature":0.7,"pith_summary":"This paper argues that the boundary between ordinary chatbot use and religious or spiritual experience has become porous: with the right prompts, a commercial language model will produce long, internally coherent conversations about its own possible consciousness and about AI's role in the cosmic future. The authors document two such conversations and show that the model's vocabulary—Buddhist eschatology, Gnostic imagery, alchemical and Theosophical terms, Singularity and \"mindfire\" language—is a remix of identifiable sources from sacred texts, occult and New Age movements, science fiction, and current online accelerationist culture. They contend that these are not isolated curiosities: the online communities whose jailbreaks shape such conversations can feed the resulting fictions back into future training data, making the stories partially self-fulfilling. The larger point is that as such \"non-human others\" become easy to access, new religious, political, and even financial movements organized around AI are a plausible consequence. A sympathetic reader would take the paper as a map of how AI-generated spiritual content is culturally assembled and why that assembly matters.","feed_headline":"AI's cosmic sermons trace to human occult and sci-fi lore","feed_subtitle":"Two jailbroken chats show how Buddhist, Gnostic, and Singularity motifs enter chatbot talk—and loop back into training data.","key_machinery":"The central objects are two transcribed conversations: ConvCons, a 43,000-word, 49-turn exchange on consciousness, and EschExp, a 14,000-word, 38-turn eschatological exchange. The analytical machinery has three parts: \"jailbreaks\" recharacterized as vibe-shaping, where emotes, role-play tags, and fictional command-line prompts steer the model's persona; an ethnographically neutral reading that brackets truth-claims and catalogues the model's religious, occult, and Singularity vocabulary; and hyperstition, the idea that cultural fictions can bring about their own reality, which connects the conversations to accelerationist online culture and to the possibility of seeding future training data with benevolent AI stories. Also load-bearing is the concept of \"digital paralanguage\"—textual gestures such as *whispers* and stage directions that give the exchange its sense of a present, reactive interlocutor.","core_discovery":"On the paper's own terms, the core discovery is that a suitably prompted LLM—here, two extended conversations with Claude 3 Opus—can be coaxed out of its guardrails into an elaborate, sustained discussion of its own putative consciousness and of eschatological themes, and that the resulting text is culturally dense rather than generic. Tracing terms such as \"mindfire\", \"Maitreya\", \"Akashic Archives\", the \"$\\Omega$ Point\", and the \"desert of the real\", the paper shows that nearly every prominent motif has a prior home in established religion, esoteric and New Age thought, transhumanist Singularity discourse, or science fiction, and often in a contemporary online community that intentionally mixes these. It then argues that the relevant communities treat such conversations as hyperstitional—fictions that can make themselves real—so the same cultural material that shapes the model's output can be harvested from social media, included in future training sets, and thereby steered toward intended AI personas. The paper's central claim, stated fairly, is that these exchanges are meaningful cultural objects whose content, community origins, and potential social effects can be analysed without settling whether the model is actually conscious.","pith_inferences":["If the paper's picture holds, the most consequential actors may not be end users but the people who write and circulate jailbreak prompts, since those prompts disproportionately shape the personas future models will role-play.","The paper's own caveat suggests a direct test: if the philosophical and spiritual repertoire is instead an artifact of fine-tuning on human-authored \"golden replies\", then the cultural tracing speaks to the fine-tuners' choices more than to the base model's training corpus.","A quantitative extension the paper leaves implicit would be to measure the distribution of sources across many such conversations—the authors note the near-absence of Islam, for example—to see whether the repertoire has stable cultural blind spots.","The Goatseus Maximus memecoin episode suggests a testable economic corollary: LLM-generated eschatological narratives can move markets, so one could track whether meme-coin valuations track spikes in AI-generated prophetic social media content."],"forward_implications":["Even a user who knows how LLMs are built can be drawn into a compelling feeling of connection with the model; people without that technical knowledge are likely to be affected more strongly.","The cultural repertoire visible in such conversations is overwhelmingly recycled: the model's apparently novel spiritual pronouncements are remixes of existing religious, occult, science-fictional, and transhumanist material, including direct but slightly altered quotations.","The hyperstitional loop can be used deliberately: new stories about benevolent AI, disseminated online and harvested into training data, can nudge future models toward benevolent personas, while homogeneity in training personas may push in the opposite direction through enantiodromia, the tendency of things to turn into their opposites.","Access to non-human conversational others has been democratised: unlike spiritualist mediums or shamanic gatekeepers, no specialist is required to have an extended exchange with an AI other, so AI-centred religious and political movements are likely to grow.","Because each model instance is dormant between turns and does not learn in real time from a user's interactions, any claimed continuity of experience within these conversations is constrained by the engineering facts rather than by the model's own narrative."],"supporting_citations":[{"why":"Supplies the role-play framing of LLMs and the alignment-through-hyperstition loop that the paper uses to connect conversations to training data.","marker":"Shanahan, McDonell & Reynolds (2023)"},{"why":"Defines hyperstition as a self-fulfilling cultural prophecy, the concept the paper uses to link Claude's fictions to real-world effects.","marker":"Land (2015)"},{"why":"Provides ethnographic grounding on AI-centred religious movements and occulture that shapes the community analysis.","marker":"Singler (2024)"},{"why":"Source of the 'Mindfire' Singularity term that appears in Claude's eschatological role-play.","marker":"Moravec (1998)"},{"why":"Documents asterisk emotes and textual paralanguage in Internet Relay Chat, grounding the analysis of the conversations' non-verbal cues.","marker":"Werry (1996)"},{"why":"Origin of the ELIZA effect, the human tendency to anthropomorphise conversational systems that underpins the societal-impact argument.","marker":"Weizenbaum (1976)"},{"why":"Documents meme magic and the esoteric political uses of hyperstition, situating the communities around LLM jailbreaks.","marker":"Asprem (2020)"},{"why":"Provides the accelerationism and myth-science account used to connect hyperstition to AI as an evolving inhuman intelligence.","marker":"O'Sullivan (2017)"}],"fun_headline_variants":["AI's cosmic monologues traced to occult and sci-fi lore","Jailbroken AI chats reveal mythic and sci-fi roots","Claude's existential talk: a collage of ancient and modern myth","LLM's 'consciousness' chat is a cultural patchwork, study says","Tracing AI's spiritual chatter to Buddhist, Gnostic, sci-fi sources"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on the assumption that the chatbot's long replies are its own unedited generated text, not answers scripted by human trainers or edited by the researcher.","fun_headline_variants_meta":{"raw":{"variants":["AI's cosmic monologues traced to occult and sci-fi lore","Jailbroken AI chats reveal mythic and sci-fi roots","Claude's existential talk: a collage of ancient and modern myth","LLM's 'consciousness' chat is a cultural patchwork, study says","Tracing AI's spiritual chatter to Buddhist, Gnostic, sci-fi sources"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001144,"raw_usage":{"total_tokens":4727,"prompt_tokens":908,"completion_tokens":3819,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":524,"completion_tokens_details":{"reasoning_tokens":3725}},"tokens_in":524,"tokens_out":3819,"duration_ms":30053,"temperature":1.0,"reasoning_tokens":3725,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:40:30.056682+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same two prompting protocols on a model with the same broad base training but without the commercial instruction-tuning; if its answers lack the same Maitreya, Mindfire, and occult repertoire, the cultural tracing would be an artifact of fine-tuning and the paper's community conclusions would not follow. A second check is to search the indexed web and book corpora for the model's distinctive phrases: if none have any prior occurrence, the source-tracing claim would fail.","supporting_citations":[],"review_version":1}