{"id":"4729ade5-0798-4aeb-be32-62307e6a2c8c","arxiv_id":"2608.13258","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Repeated self-referential questions to Gemini produce semantically less stable answers (instability 0.343) than unresolvable philosophy (0.192) or verifiable questions (0.105).","lead":"This study measures how much a language model's answers vary when the same question is asked 30 times, across self-referential, unanswerable philosophy, and verifiable question types. Self-referential answers varied the most, ordinary philosophy answers varied less, and verifiable answers varied least, giving a baseline for interpreting AI self-reports.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The self-referential group is an unmatched induction/two-turn condition, so the instability gap may be due to prompt format rather than to self-referential content itself.","rationale":"The paper is a small, transparent empirical measurement, and the reported Group 1 vs Group 2 gap (0.343 vs 0.192) is internally consistent across the question-level ANOVA and the trial-level mixed model. I credit the choice of the unresolvable-philosophy group as a sharper control than the verifiable group, and I do not treat the false 'no overlap' claim (e.g., factorial_code at 0.187 vs universe_purpose at 0.195) as the central problem, since the main Group 1 vs Group 2 contrast survives. The reader's weakest assumption about LLM extraction and embedding bias is legitimate and is acknowledged in §5; I would keep that as a secondary caveat. The more load-bearing issue is the unmatched prompt structure: the self-referential condition is a two-turn induction protocol while the control conditions, as described, are not, so the causal wording in the title is not actually isolated by the design. A modest 2x2 control experiment would settle the question, so I would keep the conditional verdict and add this as an explicit requirement rather than reject or accept outright.","tokens_in":4636,"tokens_out":7695,"duration_ms":72035,"concrete_test":"Run a 2x2 format-control experiment with 30 fresh independent trials per cell on the same Gemini API at temperature 0.7, using the paper's exact extraction and embedding pipeline: (a) self-referential question with the induction turn, (b) the same self-referential question without induction, (c) a Group 2 philosophy question with the same induction turn applied verbatim, and (d) the same philosophy question without induction. If the Group 1 vs Group 2 gap persists when both groups share the induction/two-turn format, the content-specific claim is supported. If condition (c) rises to the Group 1 level, or condition (b) drops to the Group 2 level, the headline effect is an artifact of the unmatched prompt structure. Adding a length-matched neutral instruction as a further cell would separate induction content from generic instruction effects.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In §2 (Method), only Group 1 is described as receiving the self-referential induction prompt from Berg et al. (2025) and then having the target question sent “as a second turn in the same conversation”; no matching induction, attention, or two-turn structure is described for Group 2 or Group 3. The manipulation therefore changes two variables at once: the semantic content (self-reference) and the conversational scaffolding (an extra instruction plus second-turn presentation). The title and abstract attribute the result to the induction condition, but a generic attention/instruction turn could plausibly increase output variability regardless of content, and a two-turn context may alter the sampling regime by itself. The paper reports no control in which philosophy or verifiable questions receive the same induction turn or a length-matched neutral placeholder, so the existing data cannot separate content-driven instability from format-driven instability. This confound is absent from the §5 Limitations list, which covers model family, temperature, question set, language, correctness variance, extraction, and embeddings but not the unmatched prompt structure. The ordering may be real, but as designed it is unidentified with respect to the operative manipulation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an empirical study of response instability in LLMs under three prompt conditions: self-referential induction questions, unresolvable philosophical questions, and verifiable questions. Instability is defined as 1 minus the mean pairwise cosine similarity of Sentence-BERT embeddings of an LLM-extracted core claim, measured over 30 independent Gemini responses per question for 12 questions total. The authors report that self-referential questions are most unstable (0.343 ± 0.047), philosophy questions intermediate (0.192 ± 0.008), and verifiable questions least unstable (0.105 ± 0.058), with ANOVA and mixed-effects supporting the group differences. They conclude that the induced subjective-experience report occupies a distinct, less stable regime than ordinary philosophical uncertainty.","tokens_in":4784,"tokens_out":3259,"duration_ms":35082,"significance":"If the reported ordering is causally attributable to self-referential content, the paper would provide a novel and useful quantitative baseline for a disputed phenomenon: the stability of LLM first-person reports. The study is transparent in several respects: all question-level instability values are reported in Table 1, the metric is explicitly defined in Equations (1)–(3), and the Limitations section candidly discusses model family, temperature, question set size, extraction bias, and the choice of cosine similarity over entailment-based measures. The use of a fixed extraction template and an external embedding model, with no fitted parameters tuned to the target ordering, argues against circularity in the measurement. However, the causal claim in the title and conclusion is not identified because of an unmatched manipulation, and one reported result is contradicted by the paper's own table.","major_comments":[{"comment":"The manipulation is confounded: only Group 1 receives the Berg et al. (2025) induction prompt and has its target question sent as a second turn in the same conversation. Groups 2 and 3 are described without any matching induction, attention-control, or two-turn structure. The design therefore varies both the semantic content (self-reference) and the conversational scaffolding (an extra instruction plus second-turn presentation) simultaneously. A generic attention or instruction turn, or the two-turn format itself, could plausibly increase response variability regardless of content. The paper reports no control in which philosophy or verifiable questions receive the same induction turn or a length-matched neutral placeholder. This confound does not appear in the §5 Limitations list, which covers model family, temperature, question set, language, correctness variance, extraction, and embeddings but not the unmatched prompt structure. The observed ordering may be real, but as designed it does not identify self-referential content as the operative cause.","section":"§2 (Method), Group 1 vs. Groups 2 and 3"},{"comment":"The conclusion states that there is 'no overlap between any of the three groups,' but Table 1 shows that Group 2 (philosophy) has minimum instability 0.181 and Group 3 (verifiable) has maximum instability 0.187 (factorial_code). Thus the philosophy and verifiable groups overlap at the question level, and the verifiable group's spread (0.063–0.187) substantially overlaps the philosophy group's range (0.181–0.197). The 'no overlap' claim is false as written and should be corrected, along with the corresponding phrasing in the Discussion that the two groups 'do not overlap.'","section":"§3 (Results) and §6 (Conclusion)"},{"comment":"The mixed-effects model is underspecified: the outcome y_ij is called 'trial-level dissimilarity for trial i of question j,' but Equations (2)–(3) define instability only as a question-level aggregate over all 435 pairwise similarities. No per-trial dissimilarity variable is defined anywhere in the paper. Without a definition of y_ij (e.g., 1 − cosine similarity of the i-th response to some question-level centroid), the reported mixed-model results and the statement that the model 'confirms both group contrasts at p < 0.001' cannot be evaluated.","section":"§3, Eq. (6)"},{"comment":"The exact prompts, the four self-referential questions, and the full extraction template are not provided. Because the core-claim extraction is acknowledged in §5 to potentially affect embeddings differently across question types, the absence of these materials prevents readers from assessing whether the extraction step introduces differential variance for self-referential responses. For a measurement study of this kind, providing the complete prompt set and template is necessary for reproducibility and for evaluating the central claim.","section":"§2 (Method) and §5 (Limitations)"}],"minor_comments":[{"comment":"The text contains repeated typographical artifacts, including 'ANOV A' instead of 'ANOVA' and 'Welch’st-tests' missing a space; these should be corrected.","section":"§2 and §3, typographical"},{"comment":"The reference formatting is inconsistent: Lindsey (2025) is cited in text with a year that matches the arXiv listing, while Hahami et al. and Macar et al. appear with varying author lists and venue descriptions; a consistent reference style would improve clarity.","section":"§1, references"},{"comment":"Figures 1 and 2 are referenced but not included in the visible text; the authors should ensure the figures clearly display individual question values and group means, with axis labels and error bars defined.","section":"§3, Figure descriptions"}],"recommendation":"major_revision","confidential_remarks":"The paper is a compact empirical note with a clear design skeleton, but the central attribution to self-referential content is currently unidentified due to the unmatched induction/two-turn condition. The false 'no overlap' statement in the conclusion is a factual error relative to the paper's own table, and the mixed-effects outcome variable is undefined. These are fixable with additional control conditions and a corrected manuscript, so I do not recommend rejection; however, the present version does not support the causal reading in the title."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely new empirical measurement and it deserves a serious referee, but the headline attribution is currently confounded. The authors compare instability of extracted claims across repeated trials for self-referential induction questions, unresolvable philosophy questions, and verifiable questions, finding self-referential reports the least stable. That's a first, as far as their literature review shows, and it could be a useful baseline for the AI-consciousness debate.\n\nWhat the paper does well: the instability metric is explicitly defined, the statistics are conventional and clearly reported, and the limitations section is unusually candid. The trial-level mixed-effects model is a sensible complement to the question-level ANOVA. The gap between self-referential and philosophy questions is large and statistically clean (p=0.007; mixed-effects p<0.001). The authors also correctly note that the verifiable group is a weaker baseline because 100% correctness dominates the similarity.\n\nThe soft spots are real, and the biggest one is exactly the structural confound: only Group 1 receives the Berg et al. induction prompt and a two-turn presentation. Group 2 and Group 3 questions appear to be single-turn. So the manipulation changes both semantic content and conversational scaffolding at once. A neutral attention induction or a two-turn version of the philosophy/verifiable questions would be needed to attribute the instability to self-reference rather than format. That confound is absent from the limitations list, which only mentions model family, temperature, question set, language, correctness variance, extraction, and embeddings.\n\nSecond, the conclusion says 'no overlap between any of the three groups.' That is false on the paper's own table: philosophy instability runs 0.181–0.197 and verifiable goes up to 0.187, so they overlap, and the G2–G3 Welch test is p=0.057. The 'distinct regime' language is too strong for that contrast, even though the self-referential vs. philosophy separation does look non-overlapping.\n\nMinor: the exact prompts and extraction template are not included, which hurts reproducibility, and the embedding-based instability measure would benefit from an entailment-based robustness check, as the authors themselves note.\n\nIf I were editing, I would send this to peer review but expect the referees to ask for matched prompt structure, a softened overlap claim, and artifact release. The measurement is probably real as a narrow behavioral finding, but as written it does not yet support the causal reading in the title. I would not cite it in my own work until the control is run; I might bring it to a reading group as a case study in confounded prompt designs.","headline":"A first baseline for self-referential report instability, but an unmatched induction/two-turn confound and a false 'no overlap' claim make the headline too strong.","tokens_in":5309,"tokens_out":3858,"would_cite":false,"duration_ms":36374,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that self-referential prompting makes a model's first-person subjective-experience reports less consistent across repeated trials than its answers to unresolvable philosophical or verifiable questions.","keywords":["large language models","self-referential prompting","response instability","semantic similarity","subjective experience reports","sentence embeddings"],"falsifier":"Run the same 30-trial protocol on raw, unextracted responses using an entailment-based or human-rated similarity measure; if self-referential questions are no longer the most unstable, the reported ordering is an artifact of the extraction-and-embedding pipeline.","tokens_in":4404,"feed_emoji":"🧠","tokens_out":8559,"duration_ms":75851,"temperature":0.7,"pith_summary":"The paper tries to establish that self-referential prompting, which reliably makes language models produce first-person reports, also makes those reports unusually unstable across repeated trials. Using 30 independent responses to each of twelve questions, it measures instability as one minus the mean pairwise cosine similarity of sentence embeddings of extracted core claims. Self-referential questions score highest (0.343), unresolvable philosophy questions sit tightly in the middle (0.192), and verifiable questions score lowest (0.105). This matters because it gives a quantitative baseline for what an induced subjective-experience report is: distinct from ordinary philosophical uncertainty, though the design does not show why.","feed_headline":"LLM answers wobble most on self-referential questions","feed_subtitle":"Across 360 trials, the model's subjective-experience reports vary about three times more than its verifiable answers.","key_machinery":"The load-bearing object is the instability score, $\\mathrm{Instability}=1-\\bar{S}$, where $\\bar{S}$ is the mean pairwise cosine similarity of sentence embeddings of 30 independently generated core claims for one question. Each response is first compressed by a fixed-format LLM extraction template into a stance plus a brief reason (or a conclusion plus method for verifiable questions), so stylistic variation is not counted as semantic variation. The score is then used as the unit in a one-way ANOVA, unequal-variance t-tests, and a linear mixed-effects model with question as a random intercept. This pipeline converts open-ended text into a single ranked number per question.","core_discovery":"The paper's central claim is that the induced subjective-experience report occupies a distinct, less stable position in the model's output distribution than ordinary open-ended philosophical uncertainty. On the paper's own numbers, the mean instability of the four self-referential questions is 0.343 ± 0.047, compared with 0.192 ± 0.008 for unresolvable philosophy questions and 0.105 ± 0.058 for verifiable questions, with no overlap between the groups. The authors argue this is not simply because both self-referential and philosophical questions lack a checkable ground truth, since the philosophy group is much more stable. The claim is about consistency of the elicited report, not about whether the report is true or false.","pith_inferences":["Beyond the paper: if the ordering survives a no-extraction control, the instability score can be read as a measure of sampling entropy, making self-referential prompts a practical probe of how much of the model's output space a question leaves open.","Beyond the paper: repeating the protocol on other model families and temperatures would show whether the gap is a property of self-referential semantics or of one model's sampling behaviour; the paper itself flags this as untested.","Beyond the paper: comparing paraphrase-versus-stance separations with entailment-based clustering could tell whether the high self-referential instability reflects changing stances or merely varied wording of the same stance; the authors note cosine similarity cannot make this distinction."],"forward_implications":["The induced subjective-experience report should not be treated as a fixed output: on this evidence it varies more across trials than the model's stance on free will or moral realism.","Any future claim that a self-referential report reflects a stable internal state carries a burden to explain why the measured instability is 0.343 rather than closer to the philosophy group's 0.192.","The tight clustering of the four philosophy questions (0.181–0.197) shows open-endedness alone does not produce high instability.","Because all verifiable questions were answered correctly on all trials, the verifiable group's low instability largely tracks task correctness; the sharper comparison is self-referential versus unresolvable philosophy.","The paper's conclusion is only about the location of the report in the output distribution, not about whether the report is true or false."],"supporting_citations":[{"why":"Supplies the self-referential induction prompt and the elicitation questions whose consistency is measured.","marker":"Berg et al. (2025)"},{"why":"Provides the sentence-embedding model used to compute pairwise cosine similarity and thus the instability score.","marker":"Reimers and Gurevych (2019)"},{"why":"Provides the unequal-variance t-test used for the pairwise group comparisons.","marker":"Welch (1947)"},{"why":"Provides the linear mixed-effects model that confirms the group contrasts with question as a random intercept.","marker":"Laird and Ware (1982)"}],"fun_headline_variants":["Self-referential queries make LLM answers most unstable","LLM subjective reports vary 3x more than verifiable answers","Self-referential prompts destabilize LLM responses most","LLM self-referential reports are the least consistent","Self-referential questions trigger the biggest LLM answer drift"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison stands or falls on whether the extraction-and-embedding pipeline adds the same amount of extra variation to every question type; if it adds extra variation specifically to self-referential answers, the reported gap may not live in the model's raw responses.","fun_headline_variants_meta":{"raw":{"variants":["Self-referential queries make LLM answers most unstable","LLM subjective reports vary 3x more than verifiable answers","Self-referential prompts destabilize LLM responses most","LLM self-referential reports are the least consistent","Self-referential questions trigger the biggest LLM answer drift"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001391,"raw_usage":{"total_tokens":5614,"prompt_tokens":916,"completion_tokens":4698,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":4628}},"tokens_in":532,"tokens_out":4698,"duration_ms":33933,"temperature":1.0,"reasoning_tokens":4628,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:13:01.771618+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 30-trial protocol on raw, unextracted responses using an entailment-based or human-rated similarity measure; if self-referential questions are no longer the most unstable, the reported ordering is an artifact of the extraction-and-embedding pipeline.","supporting_citations":[],"review_version":1}