{"id":"586d21e0-d26e-4ee0-98dc-a371d72b4f49","arxiv_id":"2506.11945","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Both AI researchers and the US public see AI subjective experience as a likely reality by 2100, while disagreeing on how to treat and govern such systems.","lead":"A survey of 582 AI researchers and 838 members of the US public finds that both groups think AI systems with subjective experience are more likely than not to exist by 2100, though they are divided on rights and governance. The results give policymakers and AI developers a systematic snapshot of expert and public attitudes on machine consciousness and welfare.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Construct validity is the weak point: in-survey evidence (SI Fig. 13) shows only about half of each sample treats the stipulated 'subjective experience' as required for sentience/consciousness, so the headline 2100 forecasts may not target the intended concept.","rationale":"The reader identified the same weakest assumption: respondents may not share the stipulated definition of 'subjective experience,' and responses may be ad hoc and unstable. I agree, and the paper's SI provides internal evidence that makes this concern concrete rather than speculative: only about half of each sample selected subjective experience as required for consciousness or sentience, even though the survey definition was intended to make it central to both. This is the most load-bearing concern for the central descriptive claim because it affects what quantity the headline medians actually estimate. Other concerns, such as AI-researcher sample representativeness or the fixed-probability framing artifact, are real but either acknowledged, secondary to the main headline, or less directly connected to the core claim. The paper has genuine strengths: preregistration, neutral recruitment, randomized blocks, and transparent limitations. I would therefore keep the reader's CONDITIONAL verdict rather than escalate; the proposed subgroup analysis is a feasible condition that could be satisfied with existing data once released. If the definition-consistent subgroup reproduces the headline medians, the concern is resolved; if not, the descriptive claims need substantial reinterpretation.","tokens_in":50320,"tokens_out":4157,"duration_ms":59816,"concrete_test":"Using the archived survey data (OSF, pre-registration and data to be released), split each sample into 'definition-consistent' and 'inconsistent' respondents based on the preliminary capacity task (SI Figure 13): consistent = selected 'subjective experience' as required for sentience (or consciousness); inconsistent = did not. Recompute the median 2100 probability, the 'never' probability, and the percentage agreeing that developers should implement safeguards now within each subgroup. If the definition-consistent subgroup's 2100 median differs from the inconsistent subgroup by more than 10 percentage points, or if it no longer exceeds 50%, the headline claim is an artifact of heterogeneous interpretation and must be re-reported with a definition-consistent subsample or a more robust definition manipulation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that AI researchers and the public think AI subjective experience is more likely than not by 2100 (medians 70%/60%)—requires that respondents answer about the stipulated definition: 'the ability to experience the world from a single point of view... something it is like to be that AI system.' The paper's own data cast doubt on this. In the preliminary capacity task (SI Section A.2, Figure 13), only about half of each sample selected 'subjective experience' as required for sentience (public 49%; AI researchers 55%) or consciousness (AI researchers 53%). Since the stipulated definition is supposed to make subjective experience essentially equivalent to sentience/consciousness, this is evidence that many respondents used a different folk concept even after reading the definition. The authors acknowledge the risk in Section 3.1 (folk conceptions; ad hoc responses), but the consequence is not merely a caveat: if a large minority are forecasting 'AI consciousness' in a looser sense, the reported medians and the 'more likely than not by 2100' statement are not estimates of the same quantity for the full sample. This is load-bearing because the paper's headline and policy-relevant conclusions (e.g., 68%/85% support for safeguards now) are descriptive claims about beliefs regarding the stipulated concept, not about the folk concept. The concern is testable using existing data, because the survey already contains a within-subject consistency check.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a large survey of 582 AI researchers who published in leading venues and 838 nationally representative US participants, asking about beliefs concerning AI systems with subjective experience. It measures probabilistic forecasts for when such systems might exist, views on how we could determine and recognize subjective experience, and attitudes toward the moral status, rights, responsibilities, and governance of such systems. The headline findings are that median respondents in both groups give roughly 70% (AI researchers) and 60% (public) probability that AI systems with subjective experience exist by 2100, that the public assigns a higher probability than AI researchers to the possibility that such systems will never exist (median 25% vs. 10%), and that majorities in both groups support implementing safeguards now, while views on rights, welfare protections, and whether such systems should be created are divided. The authors emphasize that the forecasts should be treated as descriptive of current attitudes rather than as predictive indicators.","tokens_in":50592,"tokens_out":4090,"duration_ms":56230,"significance":"If the results hold, this is the first large-scale comparative survey of AI researchers and the public on AI subjective experience, providing a valuable empirical foundation for policy and ethics debates. The study is preregistered, uses a large expert sample with a neutral recruitment message, applies Holm-Bonferroni corrections for multiple comparisons, and includes unusually careful and detailed limitations sections. The authors are appropriately cautious that the forecasting numbers reflect attitudes, not predictions. The central construct-validity issue—whether respondents are answering about the stipulated definition of subjective experience—is raised in the authors' own limitations but needs a direct empirical response, because it bears on the paper's main descriptive claim.","major_comments":[{"comment":"The headline forecasts (median 70% by 2100 for AI researchers, 60% for the public) are claimed to describe beliefs about the stipulated definition of subjective experience, which 'subsumes common conceptions of both consciousness and sentience.' However, the survey's own preliminary capacity task shows that only about half of each sample selected 'subjective experience' as required for sentience (public 49%, AI researchers 55%) or for consciousness (AI researchers 53%). This is direct evidence that a substantial fraction of respondents used a different folk concept even after reading the definition. The consequence is not merely a caveat: if a large minority were forecasting 'AI consciousness' in a looser sense, the reported medians and the 'more likely than not by 2100' statement are not estimates of the same quantity for the full sample. Because the survey contains within-subject data, this is testable. I request that the authors report the forecasting and governance results separately for respondents who, in the capacity task, treated the stipulated subjective experience as required for sentience/consciousness, and for those who did not. If the restricted analyses materially change the headline estimates, the abstract and executive summary should be qualified accordingly. This analysis is load-bearing because the paper's central claim and policy conclusions are meant to track the stipulated concept, not a folk variant.","section":"Section 3.1 and SI Section A.2, Figure 13"},{"comment":"The authors acknowledge that the fixed probability framing omitted a 'never' option and may have biased forecasts earlier (executive summary; Section 2.1). Yet the paper's combined aggregate forecasts—e.g., the statement in the Discussion that 'the median AI researcher and member of the US public believed there is a 10% chance... by 2030, a 50% chance... by 2050, a 90% chance... by 2100'—are built by pooling both framings and fitting distributions. If one framing is systematically biased, the pooled aggregate inherits that bias. The authors note that 'it is reasonable to take the fixed date results as more reliable than the fixed probability results,' but then present the pooled results as headline summaries. I request that the main text present fixed-date-only and fixed-probability-only aggregates as the primary summary (as in Table 2), or provide a sensitivity analysis showing that the pooled '90% by 2100' claim is robust to excluding the fixed probability framing. Without this, the pooled claims in the Discussion are not fully supported by the authors' own reliability assessment.","section":"Section 2.1 and Section 3, Figure 1(B), Table 2"}],"minor_comments":[{"comment":"The sentence 'An independent samples t-test revealed a statistically significant difference in the mean probability estimates between AI researchers in the two groups' contains a phrasing error: the comparison is between AI researchers and the public, not 'between AI researchers in the two groups.'","section":"Section 2.1"},{"comment":"The abstract states that 'Both groups perceived a need for multidisciplinary expertise to assess AI subjective experience,' but Figure 3 shows several substantial and significant between-group differences in which expertise sources are valued (e.g., AI ethics researchers, the AI system's own outputs, the public). Consider noting in the abstract or summary that the public places greater weight on several non-technical sources than AI researchers do.","section":"Abstract and Executive summary"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is well within the scope of cs.CY and is a well-executed descriptive survey. The main obstacle to acceptance is the construct-validity concern: the paper's own SI data suggest that only about half of each sample used the stipulated definition of subjective experience. This is directly addressable with the existing data, and I would want to see the restricted analyses before publication. The authors' thorough and honest limitations section helps, but the central claim needs the additional empirical support. If the restricted analyses confirm the headline findings, I would be happy to see the paper accepted after those revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is the first systematic comparison of AI researchers and the US public on AI subjective experience, and it fills a real gap. The headline finding—both groups think it is more likely than not that AI systems with subjective experience will exist by 2100 (median 70% for researchers, 60% for the public)—is new and policy-relevant. The study also adds useful data on how people think we should determine subjective experience, on moral status and rights, and on governance preferences. The design is careful: preregistered, identical questions across groups, both fixed-date and fixed-probability forecast elicitation, multiple-comparison corrections, and a transparent limitations section. The authors deserve credit for flagging the folk-conception problem and the instability of ad hoc responses.\n\nNow the soft spots. The stress-test concern about construct validity is real, but I think it is overstated as stated. The in-survey evidence from SI Figure 13 comes from a preliminary task that used a brief definition of subjective experience—\"experiences that feel like something from a single point of view\"—whereas the main forecast questions used the full, more elaborate definition. So the fact that only about half of each sample selected subjective experience as required for sentience or consciousness is suggestive, not conclusive, about whether respondents in the main block understood the target concept. Still, the authors should do the obvious sensitivity analysis: re-restrict the sample to respondents who did map subjective experience onto sentience/consciousness and see if the headline forecasts change. That is a cheap, testable fix, and the current manuscript lets the concern linger.\n\nTwo other caveats, both minor. The AI researcher sample has a low response rate and skews academic, male, and Western; the authors note this but it does limit generalization. And the distribution fitting deviated from the preregistered analysis; the authors explain why, but preregistration deviations should be highlighted more prominently than they are. Data and code are not yet available, which is a genuine condition on verification.\n\nThe paper is not the last word on any of these questions, but it is a serious empirical contribution. If I were editor, I would send it to peer review with a request for the sensitivity analysis and public data/code. The descriptive claims about attitudes are supportable; the specific probability values should be treated as approximate.\n\nReading group: yes. It will generate a good discussion about measurement validity in AI attitudes research.","headline":"A solid, preregistered survey that gives the first direct comparison of AI researchers and the US public on AI subjective experience; the headline forecasts are broadly credible, but the construct validity question deserves a sensitivity analysis before the specific numbers are taken as precise.","tokens_in":51115,"tokens_out":2326,"would_cite":true,"duration_ms":133127,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Two large surveys, one of 582 AI researchers and one of 838 US adults, find that both groups regard AI systems with subjective experience as more likely than not to exist by 2100 (median 70% and 60%) and that majorities in both groups…","keywords":["AI subjective experience","public opinion survey","expert survey","consciousness forecasting","AI governance","moral status of AI","sentient AI","safeguards"],"falsifier":"Two concrete tests would settle the matter: re-contacting the same respondents weeks later to see whether the median 2100 probabilities hold, and re-running the survey with alternative definitions — one emphasizing valenced experience (pleasure and pain, which the paper's own auxiliary results show shifts timelines later) and one that adds a \"never\" option to the fixed-probability framing. If the medians fall below 50% under either change, the headline \"more likely than not by 2100\" result is a framing artifact rather than a stable belief.","tokens_in":50130,"feed_emoji":"🤖","tokens_out":13776,"duration_ms":152311,"temperature":0.7,"pith_summary":"This paper reports a survey of 582 AI researchers and 838 US adults asking whether and when AI systems might have subjective experience — an internal life from a single point of view, such that there is \"something it is like\" to be the system — and how such systems should be treated and governed. The central finding is that both groups regard such systems as more likely than not to exist by 2100, with median probabilities of 70% among AI researchers and 60% among the public, and no statistically significant difference between the groups' timelines. The public is markedly more open to the possibility that such systems will never exist (median 25%, versus 10% for researchers) and is more cautious about building them. Yet majorities of both groups agree that AI developers should implement safeguards now (68% of researchers, 85% of the public) and that such systems, if created, should be held accountable and behave ethically, even as views on rights and moral status are divided. If the results are right, there is broader agreement than is often assumed — both on the realistic possibility of machine minds this century and on precautionary action in the meantime.","feed_headline":"Experts and the public expect AI minds by 2100, survey finds","feed_subtitle":"Median odds: 70% (AI researchers) and 60% (US public), with majorities in both groups backing developer safeguards now.","key_machinery":"The load-bearing instrument is the survey battery itself, anchored by a stipulated definition: subjective experience is \"the ability to experience the world from a single point of view, including experiences such as perceiving... and feeling,\" so that an AI system with it would have \"something it is like\" to be that system — a definition that subsumes common conceptions of consciousness and sentience. Forecasts were elicited in two framings, fixed dates (2024, 2034, 2100) and fixed probabilities (10%, 50%, 90%), and each respondent's answers were encoded as a continuous probability distribution by a sequential fitting procedure — metalog, then a quantile-parameterised distribution, then a three-component skew-normal mixture, then a shifted log-normal — which allows group-level aggregation by mean and by median. A separate question asking for the probability that such systems will never exist does independent work: it exposes a public-researcher split (median 25% versus 10%) that the timeline questions alone would have concealed. The conjoint safeguard experiment, in which explicitly noting the absence of safeguards lowered support for a hypothetical AI project, serves as a behavioral check on the stated governance attitudes.","core_discovery":"The paper's claim, stated on its own terms, is that technical AI researchers and the general US public hold similar, policy-relevant beliefs about AI subjective experience. Both groups think current systems almost certainly lack it (median probability 1% and 5% for 2024), put a 25% and 30% median chance on it within a decade, and a 70% and 60% median chance by 2100; fitting each respondent's forecast into a continuous probability distribution puts the median 50% likelihood at 2050 for both groups. Both groups think we would be only somewhat more likely than not to recognize such a system if it existed (median 60%), believe that between \"somewhat\" and \"moderately\" confident would be enough to grant some moral consideration, and rate technical AI researchers, neuroscientists and psychologists, and AI ethics researchers as the most important voices in any determination. On treatment and governance, majorities in both groups agree that such systems should be held accountable for their actions and should behave well, and that developers should start implementing safeguards now; support for protecting their welfare is real but far weaker than support for protecting animals or the environment, and socio-political rights are rejected by pluralities. The paper reads this pattern as evidence that both groups take machine subjective experience seriously as a live possibility this century, while remaining uncertain about the timeline and divided on the response.","pith_inferences":["The 70% and 60% medians likely overstate stable belief: the paper's own auxiliary results show that only about half of each sample regards subjective experience as required for consciousness or sentience, and the authors report strong framing sensitivity, so a deliberative or repeated elicitation would probably produce lower and more divided numbers.","The public's higher \"never\" probability (25% versus 10%) may foreshadow political friction: as AI systems become more human-like, the public's greater tendency to see inner life in them could push for stronger precaution than the technical community, which mostly thinks advanced performance needs no experience, is willing to accept.","Because both groups would act on only moderate confidence, the real governance question is not \"are they conscious?\" but how institutions should behave under diagnostic uncertainty — a question the paper documents but does not itself answer."],"forward_implications":["Policymakers can build on a shared majority premise: both experts and the public want AI developers to implement safeguards against risks from subjective AI now, years before such a system is plausibly claimed to exist.","Future AI-forecasting surveys should include a \"never\" probability question, because the timeline questions alone concealed a significant public-researcher difference in skepticism.","Governance debates should expect a public that is more precautionary than the technical community — more supportive of bans, of regulation now, and of never building such systems — while researchers are readier to build.","Any credible process for determining AI subjective experience will need to combine technical AI, neuroscience, psychology, and AI-ethics expertise; both groups rated policymakers' views as least important.","Even if subjective experience in AI were established, welfare protection would face an uphill political path: support sits far below that for animals and the environment and close to that for business corporations."],"supporting_citations":[{"why":"Supplies the \"something it is like\" phrasing that anchors the survey's stipulated definition of subjective experience.","marker":"Nagel 1974"},{"why":"Demonstrated that fixed-date versus fixed-probability framing changes AI forecasts, motivating the survey's dual-framing design.","marker":"Grace et al. 2018"},{"why":"Earlier forecasting survey of machine-learning researchers whose framing methodology the present design directly adapts.","marker":"Zhang et al. 2022"},{"why":"Provided the sampling frame of AI researchers (recent authors at leading machine-learning venues) and the neutral-recruitment template.","marker":"Grace et al. 2024"},{"why":"The prior AIMS survey whose much earlier sentient-AI forecasts this study contrasts with its own 2050 and 2100 timelines.","marker":"Pauketat et al. 2023"},{"why":"Prior public-attitude work on sentient AI establishing the rights-and-protections pattern this study replicates and extends.","marker":"Anthis et al. 2024"},{"why":"Proposed indicator list for AI consciousness, the reference point for asking whether experts think such systems could be recognized.","marker":"Butlin et al. 2023"}],"fun_headline_variants":["Researchers and public: 70% and 60% median odds of AI feelings by 2100","AI experts and the public: 60-70% chance of conscious AI by 2100","Both groups agree: AI subjective experience likely by 2100","Survey: Researchers and public both expect AI minds this century","Public and researchers: 60-70% odds AI gains subjective experience by 2100"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey results stand or fall on the assumption that respondents understood \"subjective experience\" the way the paper defines it and that their answers reflect stable beliefs rather than snap judgments about an abstract and unfamiliar topic; the paper itself flags this, noting that responses may have been formed ad hoc during the survey and could be unstable or sensitive to question framing.","fun_headline_variants_meta":{"raw":{"variants":["Researchers and public: 70% and 60% median odds of AI feelings by 2100","AI experts and the public: 60-70% chance of conscious AI by 2100","Both groups agree: AI subjective experience likely by 2100","Survey: Researchers and public both expect AI minds this century","Public and researchers: 60-70% odds AI gains subjective experience by 2100"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001644,"raw_usage":{"total_tokens":6612,"prompt_tokens":1108,"completion_tokens":5504,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":724,"completion_tokens_details":{"reasoning_tokens":5398}},"tokens_in":724,"tokens_out":5504,"duration_ms":45447,"temperature":1.0,"reasoning_tokens":5398,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:59:24.622631+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Two concrete tests would settle the matter: re-contacting the same respondents weeks later to see whether the median 2100 probabilities hold, and re-running the survey with alternative definitions — one emphasizing valenced experience (pleasure and pain, which the paper's own auxiliary results show shifts timelines later) and one that adds a \"never\" option to the fixed-probability framing. If the medians fall below 50% under either change, the headline \"more likely than not by 2100\" result is a framing artifact rather than a stable belief.","supporting_citations":[],"review_version":1}