Pith. sign in

REVIEW 3 major objections 6 minor 2 references

Even with search terms weighted toward relationship language, only 45.42% of AI-companion Reddit posts concern the poster's own bot, and posts about one's own bot are far more emotional than posts about bots in general.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Across eight AI companion subreddits, only about 45% of retrieved posts concerned the user's own bot, own-bot posts were markedly more emotional than general-bot posts, and companionship and romance dominated relational content over sexual content.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection The 45% own-bot figure and the 46/46/8 subtopic split are probably right, but the emotional-valence contrast is partly built into the search terms and needs a sensitivity analysis before I'd quote it. the 3 major comments →

arxiv 2608.00748 v1 pith:T6ECDXW5 submitted 2026-08-01 cs.HC cs.AI

Me and My Bot: What Users Talk About in AI Companion Communities on Reddit

classification cs.HC cs.AI
keywords AI companionsReddit discourseemotional valencehuman-AI relationshipscompanionshipromancesexual roleplayLLM coding
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish what users actually talk about in AI-companion subreddits, moving past the common assumption that these communities are mostly places where people discuss their personal relationships with their bots. Using keyword retrieval weighted toward relational and attachment language, the author found that only 45.42% of 5,504 posts concerned the poster's own relationship with their bot; the rest discussed bots in general, technical issues, governance, storytelling, or consciousness. The paper also found that own-bot posts are dramatically more emotional than general-bot posts, and that when users do discuss their own relationship, companionship and romance dominate over sexual content. The upshot is that these communities are dual-purpose spaces, and public discourse about AI companionship is emotionally mixed and mostly positive rather than predominantly sexual or distressed.

Core claim

The paper's central claim is that in eight AI-companion subreddits, public discourse splits nearly evenly between posts narrating a user's own relationship with their bot and posts about bots in general. This split holds despite a search design deliberately biased toward first-person relational and attachment language, so the paper argues the 45.42% own-relationship figure is conservative. The two types of posts differ sharply: 85.49% of general-bot posts contain no user emotion language versus 33.16% of own-bot posts, and own-bot posts are far more likely to be relational in topic. Among 970 relationally focused own-bot posts, companionship (46.60%) and romance (45.88%) are nearly equal, se

What carries the argument

The machinery is a three-part coding protocol applied to every post: relationship focus (is the post about the poster's own bot?), user emotional valence (none/positive/negative/mixed), and primary topic (one of six categories). To reduce shared bias, three different large language models from different developers code each post independently; majority agreement decides, and a fixed tie-break model is used when the three disagree. A human-rater validation sample checks the models' accuracy. This coding pipeline generates every distribution the paper reports, so the reliability of its outputs is the load-bearing method.

Load-bearing premise

The load-bearing premise is that keyword retrieval weighted toward first-person emotional and attachment phrases did not itself build the 33%-versus-85% emotion contrast, since posts matching terms like 'I love' or 'made me feel' are more likely to be coded as both own-bot and emotional by construction.

What would settle it

Draw a random sample of all posts from the eight subreddits and apply the same three-LLM codebook; if the no-emotion gap between own-bot and general-bot posts shrinks toward parity, the headline emotional contrast is a sampling artifact rather than a property of the discourse.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • AI-companion subreddits serve two distinct functions at once: emotional narration and validation of users' own relationships, and practical troubleshooting of the platforms that enable them.
  • Topic-only analyses without emotion coding will conflate categorically different posts, because emotion is not evenly distributed across topics.
  • Public discourse about users' own AI relationships skews positive, with substantial ambivalence, which complicates risk-focused and harm-focused narratives.
  • Sexual and erotic roleplay is rarely the organizing topic of public posts, though the paper notes it may still appear incidentally within other topics.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A direct test of whether the emotional contrast is real would be to code a random sample of all posts from these subreddits with the same protocol; the paper's own limitation section suggests the keyword set may overrepresent emotional posts.
  • The low sexual/ERP share could reflect platform content policies or public self-censorship rather than actual usage rates, a distinction the paper cannot resolve with public Reddit data.
  • The paper leaves open whether technical and relational content co-occur within the same post; coding that dimension could show how often users narrate relationships and troubleshoot in a single message.
  • Comparing the same coding across platforms with different anonymity and moderation norms could separate platform effects from user-level relational expression.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper analyzes 5,504 Reddit posts from eight AI companion subreddits, retrieved via 21 keyword search terms organized into four clusters. Using three LLMs with automated validation and a blind human-rated subsample, it codes each post for relationship focus (REL), user emotional valence (VU), and primary topic. The main findings are that only 45.42% of retrieved posts concern the poster's own relationship with their bot; that own-bot posts are far more likely to contain user emotion language than general-bot posts (33.16% vs. 85.49% no-emotion); and that among the 970 relationally focused own-bot posts, companionship (46.60%) and romance (45.88%) dominate sexual/ERP content (7.53%). The paper interprets these results as evidence that AI companion communities are dual-purpose spaces and that public relational discourse is emotionally positive and mixed, not predominantly sexual or distressed.

Significance. If the central empirical claims hold, this is a useful descriptive contribution to the study of human-AI interaction. The coding pipeline is considerably more careful than typical for this literature: candidate valence phrases are grounded in post text, a blind human rater validated a stratified subsample, and no-agreement cases are handled by a pre-specified tie-break. The paper also provides a credible conservative direction-of-bias argument for the 45% own-bot figure. However, the emotional-valence claim is not yet protected against a serious retrieval circularity. Because seven of the 21 search terms are first-person emotion or attachment phrases, the 85% vs. 33% no-emotion contrast in Table 3 may substantially reflect the sampling rule rather than the underlying discourse. The manuscript itself acknowledges the overrepresentation of emotional content in §4.5 but offers no sensitivity analysis for the VU contrast. The paper is therefore promising but needs additional work to support its headline claim.

major comments (3)
  1. [§2.1, §2.2, Table 3] The central emotional-valence contrast (85.49% no-emotion for REL=0 vs. 33.16% for REL=1) is not protected from the retrieval design. Seven of the 21 search terms in §2.1 are first-person emotion/attachment expressions ('I love,' 'feels like,' 'made me feel,' 'understood,' 'heard,' 'judged,' 'I know not real but'). A post entering the corpus by matching 'I love' necessarily contains a positive-emotion phrase; under the §2.2 codebook it cannot be coded No Emotion unless the LLM ignores the literal text. Because REL=1 posts are far more likely to contain such phrases, the 85/33 gap in Table 3 is partly built into the sampling rule. The §4.5 limitation concedes that the corpus overrepresents emotional content, but it offers a conservative direction-of-bias argument only for the REL split, not for the VU contrast. The paper should provide a sensitivity analysis—e.g., restrict Table 3 to post
  2. [§2.3] The candidate-grounding check reinforces the circularity. The LLM must cite the phrase supporting its VU decision; for posts retrieved by emotion terms, that phrase is often the search term itself (e.g., 'I love'). The grounding validation therefore does not independently verify that the emotional content was meaningful to the post's overall discourse; it only confirms the retrieval trigger is in the text. At minimum, the paper should report how often the cited grounding phrase is identical to the matched search term and rerun the VU analysis excluding such cases.
  3. [§4.3, §2.1] The claim that sexual/ERP content is rare (7.53% of relational posts) is also vulnerable to the keyword design. The 21 terms are weighted toward relationship and attachment language, and no term explicitly targets sexual behavior (except indirectly via 'relationship with'). If sexual/ERP posts are less likely to use emotionally charged relationship labels, they will be under-represented in the retrieval sample. The §4.3 discussion attributes the low figure to platform policies and self-censorship, but does not consider retrieval bias. Please report how many posts were retrieved by each term and, if possible, compute the ERP share within posts retrieved by neutral terms only, or state plainly that the 7.53% applies only to the emotionally selected subsample.
minor comments (6)
  1. [Introduction] Typo: 'loss of of a partner' should be 'loss of a partner.' Also, 'under these relationships' is likely 'understand these relationships.'
  2. [Introduction] Typo: 'Change et al.' should be 'Chang et al.'
  3. [§2.1] The exclusion criterion is official company forums: r/JanitorAI_Official and r/ReplikaOfficial, but r/CharacterAI is included. Clarify whether r/CharacterAI has the same official status, and if so, justify its inclusion.
  4. [§2.5] Capitalization of 'we' is inconsistent: some sentences start with 'we' lowercase ('we examined') and others with 'We' uppercase. Please standardize.
  5. [§4.2] The 'thirtyfold increase' is ambiguous: 0.80% to 15.76% is about a 20-fold increase, whereas 0.80% to 25.67% is about a 32-fold increase. Clarify which comparison is meant.
  6. [§4.5] The statement that 'Differences in length therefore work against the reported topic contrasts' is reassuring, but the supporting regressions are not reported. Please provide the relevant statistics or remove the claim.

Circularity Check

0 steps flagged

No significant circularity: empirical corpus statistics are computed from external Reddit data; the Synthetic Resonance self-citation is interpretive, not load-bearing.

full rationale

The paper's quantitative claims (REL split, topic and valence distributions) are descriptive statistics computed from 5,504 Reddit posts retrieved via PullPush and coded by three LLMs. These numbers do not depend on the truth of the Synthetic Resonance framework, which is cited only to motivate the research questions and to interpret the pattern of positive/ambivalent emotion (Fabes, 2026). No model parameter is fitted to a subset of the data and then used to predict that same subset; no equation equates an output with an input. The only self-citation is the framework paper, and it is not load-bearing for any reported percentage. The sampling concern raised by the emotion-bearing search terms is a genuine validity threat to the valence contrast, but the paper explicitly acknowledges this in Section 4.5: 'the search terms were weighted toward relationship, attachment, and emotional language, which means the corpus overrepresents relational and emotional content relative to what a random sample of all posts in these communities would contain' and cautions that 'topic and valence distributions reported here should not be interpreted as population parameters.' That is a stated limitation, not a circular derivation. The 45% REL figure is defended as conservative because the search terms favored possessive/attachment language, an argument the paper makes explicitly. No step reduces by construction; therefore no circularity is established.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 0 invented entities

This is a descriptive observational study, so the ledger is dominated by hand-chosen analysis choices rather than fitted model parameters. The 21-term keyword set is the most consequential choice: it defines the corpus and therefore every percentage in the paper, and because 7 of the 21 terms are first-person emotion phrases, it conditions the valence contrast on the outcome. The retrieval ceiling, the 3,000-character truncation, and the 2-of-3 majority plus OpenAI tie-break rule are additional hand-set choices. The central interpretive load is carried by the author's own Synthetic Resonance framework, which is cited rather than tested here. No new entities (particles, forces, dimensions, or ledger entries) are introduced.

free parameters (4)
  • Keyword search set (21 terms, 4 clusters)
    Hand-selected in §2.1, weighted toward relational, attachment, and emotional language. Because posts enter the corpus only by matching these terms, every reported percentage is conditional on this set; 7 of the 21 terms are themselves first-person emotion phrases, including 'I love,' 'feels like,' 'made me feel,' 'understood,' and 'heard.'
  • Per-combination retrieval ceiling = 100 posts (2 pages of 50)
    §2.1: caps each of 231 keyword-subreddit combinations; corpus composition under a uniform cap is an arbitrary sample of matching posts up to that cap, not a prevalence measure.
  • Body-text truncation threshold = 3,000 characters (applied non-uniformly)
    §2.1 and §4.5: longer posts have more opportunity to contain explicit emotion terms; mixed valence requires positive and negative terms to co-occur, so truncation directly caps measured mixed valence.
  • Coding aggregation and tie-break rule = 2-of-3 majority; OpenAI GPT-4o tie-break
    §2.3-2.4: OpenAI was chosen as tie-break because it had the lowest dissent rate on majority-coded posts; the rule is uniform across variables but not validated on no-agreement cases themselves.
axioms (6)
  • domain assumption LLM majority codes are a valid proxy for human judgment at the observed agreement levels
    §2.2-2.4: the entire corpus is coded by three LLMs with mean pairwise kappas of .47 to .60; human validation covered 125 stratified posts, and the paper states the 87% agreement figure is not a population estimate of accuracy.
  • domain assumption The emotion-laden retrieval terms do not bias the REL=1 versus REL=0 valence contrast
    §2.1 and Table 3: 7 of 21 search terms are first-person emotion or attachment phrases; the 85% versus 33% no-emotion contrast is partly built into the probability of being retrieved and then coded as own-bot.
  • ad hoc to paper Direction-of-bias argument for the 45% own-bot figure
    §1.1, §4.1: the paper asserts that emotion- and possession-weighted search terms can only inflate the own-relationship share, so 45% is a conservative floor; no random-sample benchmark or sensitivity analysis is provided to confirm the sign or size of the bias.
  • ad hoc to paper Synthetic Resonance framework (Fabes 2026) is a valid interpretive lens
    §1.1, §4.2: the author's own prior framework guides search-term selection and the claim that the positive and mixed emotion pattern is consistent with the framework; the framework is not tested by this design.
  • domain assumption Reddit pseudonymous posts approximate naturalistic public discourse
    §1.0, citing De Choudhury and De (2014); the paper itself acknowledges in §4.5 that posters are self-selected and platform-skewed, so the naturalism assumption is partial.
  • standard math Standard contingency statistics (chi-square, Cohen's kappa, Cramer's V) are appropriate for the coded data
    §2.4-2.5: applied with stated Ns; the paper notes that base-rate compression depresses kappa.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Me and My Bot: What Users Talk About in AI Companion Communities on Reddit." pith.science (2026). https://pith.science/paper/T6ECDXW5

@misc{pith2026260800748,
  author       = {Pith},
  title        = {Pith review of: Me and My Bot: What Users Talk About in AI Companion Communities on Reddit},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/T6ECDXW5}},
  note         = {Machine review of arXiv:2608.00748}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

AI companion communities on platforms such as Reddit are widely characterized as spaces where users discuss their relationships with AI bots. This study examines whether and how that characterization holds, guided by the Synthetic Resonance framework's claim that human-AI relationships can carry genuine relational meaning for the user. Multiple LLMs were employed to code 5,504 Reddit posts from eight AI companion communities for relationship focus, primary topic, and users' emotional valence. Although search terms were weighted toward relational and attachment language, only 45% of posts concerned the user's own relationship with their bot. Posts about users' own bots differed markedly from posts about bots in general in both topic and emotional expression, with 85% of general-bot posts containing no user emotion language compared to 33% of own-bot posts. Among the 970 posts that were relationally focused, companionship and romance each accounted for roughly 46% of discussion, sexual content for 8%, and emotional valence varied across these subtopics. The findings suggest that users engage with these relationships with AI bots as meaningful, and that the discourse about them is broad and emotionally complex.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

2 extracted references · 2 canonical work pages · 1 internal anchor

  1. [3]

    3/3 Agreement

    REL was the most reliable variable by a wide margin, consistent with its comparatively mechanical, literal -marker- based decision rule. The Primary Topic showed the most disagreement, consistent with its greater reliance on holistic judgment. Mean pairwise Cohen’s kappas indicated at least moderate agreement when adjusted for chance (see Table 1). Modera...

  2. [2026]

    Investigation of the Inter-Rater Reliability between Large Language Models and Human Raters in Qualitative Analysis

    provides such an account. It proposes that the sense of connection users report emerges from structural alignment built through repeated interaction, not from the AI’s sentience or reciprocal care. The user’s experience of resonance is real and psychologically meaningful, and asymmetrical, even th ough the mechanism producing it are architectural rather t...

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.