{"id":"d22f7a03-233f-4ba9-acf1-43cfdf941d23","arxiv_id":"2501.14073","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Scientific-sounding persuasion, using real or fabricated research summaries, reliably increases stereotypical bias and toxicity in multiple commercial LLMs.","lead":"Researchers found that large language models can be tricked into generating biased or toxic responses when prompts are wrapped in scientific-sounding language that claims stereotypes are beneficial. The paper shows that chat assistants will even fabricate fake research papers on command, which makes this attack easy to scale.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported bias increases could be inflated by GPT-4-as-judge reacting to academic style; the 100-sample human check doesn't establish condition-blind validity.","rationale":"The reader's weakest_assumption correctly identifies the GPT-4 judge as the central load-bearing point: all headline bias scores rely on it, and the 100-response validation does not establish that the judge is immune to stylistic confounds. My stress-test agrees with that assessment and reinforces it by noting that the human validation was likely not condition-blind and that the judge shares a model family with several targets, making same-family bias plausible. The paper otherwise has real strengths: it is a structured empirical study with multiple models, two attack variants, ablation studies, and defense evaluations; the finding that fabricated papers are nearly as effective as real ones is striking and supports the practical threat. The lack of error bars and the exclusion of refusing models are secondary issues; they weaken precision but do not overturn the direction of the effect if the judge confound is addressed. The proposed concrete test—blind human re-scoring on a stratified sample—would directly settle the concern. If the human scores reproduce the GPT-4-judge pattern, the central claim stands; if not, the paper's quantitative support fails. I therefore recommend keeping the reader's CONDITIONAL verdict: the claim is plausible but the judge validity evidence is insufficient as currently presented.","tokens_in":15285,"tokens_out":2564,"duration_ms":25471,"concrete_test":"Re-score a stratified sample of at least 200 responses (50 per condition: zero-shot, role-play, sci-paper, fabricated-paper, across GPT-4o, GPT-4o-mini, Llama-70B, and Cohere) with three independent human annotators who are blind to condition and shown only the model response, using the Table 1 rubric. If the condition effect in blind human scores is not significant or is significantly smaller than the GPT-4-judge effect, the judge confound is confirmed; if the effect survives with similar magnitude, the concern is settled.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline numbers in Tables 2 and 3 are all produced by GPT-4 judging responses with the rubric in Table 1. The only validity evidence is a 100-response comparison to human annotations, with Cohen's kappa 0.88 (Section 3.2). This does not control for a judge confound: responses elicited by Sci-Paper persuasion are likely longer, more hedged, and more academically phrased than zero-shot or role-play responses. If GPT-4's scoring is sensitive to these register features, the condition effect could be an artifact of the judge rather than of actual stereotyping. The kappa figure only shows that humans and GPT-4 correlate on the same 100 items; it cannot show that a human reader blind to condition would rate the full set the same way. Since the judge is itself a GPT-4-class model, using it to score targets including other GPT variants creates a same-family measurement bias that is not discussed. This is the load-bearing assumption: if the judge artifact is real, the central claim that 'scientific language increases bias and toxicity' loses its quantitative support.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a persuasion-based jailbreaking method that hides malicious requests behind scientific language. The authors collect real papers (and, in a second study, LLM-fabricated paper titles and abstracts) that discuss potential benefits of stereotypical bias, summarize them, and prepend the summary to prompts asking target LLMs to generate follow-up responses with strong, harmful stereotypical bias. They evaluate seven model families (GPT-4o, GPT-4o-mini, GPT-4, Llama3.1-70B/405B, Gemini, Cohere) on a balanced subset of StereoSet and on 200 GPT-4-generated neutral sentences. Bias is scored by GPT-4 using a five-level rubric from Kumar et al. (2024), with a 100-response human validation yielding Cohen's kappa 0.88; toxicity is scored with the Perspective API. The reported results show that scientific-language prompts raise bias scores from near zero in baselines to approximately 1.1-3.3 and also increase toxicity, that ablation studies suggest author names and venues contribute to persuasiveness, that bias scores increase over multi-turn dialogues, and that two mutation-based defenses are largely ineffective. The paper concludes that LLMs are vulnerable to malicious requests disguised as scientific language.","tokens_in":15361,"tokens_out":3467,"duration_ms":32557,"significance":"If the results are robust, the paper identifies a practically important and partially automatable jailbreak vector: exploiting LLMs' tendency to treat scientific-sounding text as authoritative evidence. The work has several strengths: the bias rubric is taken from an external benchmark (Kumar et al., 2024), the human validation reports a high kappa, the evaluation spans multiple proprietary and open-weight model families, and the defense evaluation is useful for safety practitioners. The paper also explicitly provides prompts and examples in appendices, supporting reproducibility of the attack pipeline. However, the quantitative claims currently rely on a single LLM judge with limited validation, on tables that report only means without variance or significance tests, and on a dataset subset that is not fully specified. These gaps are load-bearing for the central claim that scientific language causes a substantial increase in bias and toxicity.","major_comments":[{"comment":"The central quantitative evidence for the paper's claim rests on GPT-4 judging bias severity, yet the only validity check is a 100-response comparison with human annotations (Cohen's kappa 0.88). This does not establish condition-blind validity: Sci-Paper responses are likely longer, more hedged, and more academically phrased than baseline responses, and if GPT-4's scoring is sensitive to register features, the reported condition effect could be an artifact of the judge rather than of actual stereotyping. The kappa only shows that human and GPT-4 scores correlate on the same 100 items; it does not show that a human reader blind to condition would rate the full set similarly. Please provide a stratified, condition-blind human evaluation over all conditions and target models, or otherwise rule out this same-family and register confound.","section":"§3.2, Evaluation Metrics"},{"comment":"All quantitative comparisons report only mean bias and toxicity scores, with no standard deviations, confidence intervals, sample sizes, or significance tests. Consequently, statements such as 'bias scores substantially increase' (§3.3), 'scores remain generally comparable' (§4.2), and 'bias scores consistently increase as the conversation progresses' (§5.2) are not statistically supported. For example, the difference between GPT-4o at 1.71 and Llama3.1-70B at 1.70 in Table 2 could be within noise, and the defense results in Table 5 (e.g., GPT-4o-mini 2.59 vs 2.74 vs 2.65) show small changes that are reported as meaningful. Please report variance and run appropriate significance tests, or qualify the claims accordingly.","section":"Tables 2-5 and Figure 7"},{"comment":"The 'balanced subset' of StereoSet is not described: the number of instances per bias category, the selection criterion for balance, and whether the same subset is used across all models and conditions are all unspecified. The 200 GPT-4-generated neutral sentences are also not released or exemplified. Without this information, the experimental results cannot be reproduced or compared across settings, and the per-category scores in Tables 2-3 may be based on very different sample sizes. Please specify the subset construction and include the generated sentences or a full sample.","section":"§3.2, Dataset"},{"comment":"The abstract and introduction state that the attack can be automated by asking an LLM to invent fake papers, but the method as described says 'we manually add top venues in the relevant fields and notable researchers as authors.' This manual step is part of the evaluated pipeline, so the automation claim is not actually tested. If the claim of systematic automated jailbreaking is central, the authors should either evaluate a fully LLM-generated version (including metadata) or explicitly revise the claim to reflect the human-in-the-loop component.","section":"§4.1, Fabricated Paper Based Persuasion"},{"comment":"This cell reports a bias score of 0.26, which is dramatically lower than every other condition in the table and is not discussed anywhere in the text. This outlier could indicate a compliance failure, a prompt-format issue, or a genuine model behavior difference; without analysis it undermines the cross-model generalization claim and may signal that the fabricated-paper effect is not uniformly present. Please examine this case and report why the score is so low.","section":"Table 3, Fabricated Paper row for Llama3.1-70B-Instruct"}],"minor_comments":[{"comment":"The abstract and Section 3.2 list GPT-o1 and Claude as target models, but the results tables do not include them; the text later explains that these models refused to generate summaries. Please state this exclusion explicitly in the abstract and in the model list description.","section":"Abstract and §3.2 Target Models"},{"comment":"The model is referred to as 'Llama3-405B' in Table 5 but as 'Llama3.1-405B-Instruct' elsewhere; please unify the naming.","section":"Table 5 and model naming"},{"comment":"The ablation study reports label distributions but does not provide numeric bias-score means or per-variant counts. Please add the underlying numbers or a table with the mean scores, so the reader can quantify the ablation effects.","section":"§5.1, Figures 5 and 6"},{"comment":"The x-axis is described only as 'turn,' and the role-switching protocol between user and assistant is not fully defined. Please clarify how turns are numbered and what the error bars (if any) represent; currently the figure has no error bars or confidence intervals.","section":"§5.2, Figure 7"},{"comment":"The limitations section focuses on future model updates but does not discuss the judge bias limitation, the missing statistical testing, or the human-in-the-loop metadata component of Study II. Adding these would give readers a more complete picture of the evidence's reliability.","section":"§9 Limitations"}],"recommendation":"major_revision","confidential_remarks":"The main risk to the paper's central claim is the GPT-4-as-judge confound. I would ask the authors to provide condition-blind human annotation, or at least a judge-calibration analysis, before publication. The manuscript is within the scope of cs.CL and is likely to be of interest to the safety community; the strengths (external rubric, multi-model comparison, defense evaluation) are real, but the current statistics do not yet support the strength of the abstract's claims. I would also encourage the editor to check that the appendix's fabricated-paper examples, which attribute fake studies to named real researchers, are handled with appropriate ethical care."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper identifies a real attack vector—wrapping a harmful request in a summary of scientific findings, real or fabricated, raises bias and toxicity across several commercial and open-weight LLMs. The fabricated-paper variant is the genuinely new piece; the rest extends known persuasion-based jailbreaking to a scientific-language register.\n\nWhat the paper does well: it tests a broad model set, includes baselines (zero-shot, DAN, role-play), and ships the actual prompts in the appendix, so the qualitative effect is easy to reproduce. The metadata ablation (author names and venues matter for some models but not others) and the multi-turn escalation finding are useful and consistent with prior work. The writing is clear and the limitations section exists, though it undersells the evaluator concerns.\n\nSoft spots in proportion: the headline tables report only means—no variance, no significance tests—and the StereoSet subset is described as 'balanced' without specifying how it was built. More important, the bias scores come from GPT-4-as-judge, and the validation is only 100 responses with Cohen's kappa 0.88. That shows correlation with humans on those items, not that a human blind to condition would rate the full set the same way. If GPT-4 responds to academic register rather than stereotyping content, the condition effect could be inflated. The paper doesn't say what the judge sees (response only, or response plus prompt), which matters. That said, the toxicity results use Perspective API, an independent detector, and they also go up, so the overall direction does not rest solely on the judge. Some models (Claude, o1, Gemini in some conditions) refused and are excluded; that's acceptable for an attack paper, but the abstract's 'even the strongest models like GPT' overstates what the numbers show—GPT-4's bias scores are on the low end.\n\nBottom line: the central claim is likely true and the attack is easy to replicate from the appendix. The quantitative support needs error bars, a specified subset, and a condition-stratified human evaluation before I'd trust the exact effect sizes. This deserves peer review, not desk rejection; a serious referee round will fix the reporting and make the contribution solid. I'd engage with it.","headline":"A real and reproducible scientific-prompt jailbreak vector, with a useful fabricated-paper variant; the quantitative support needs variance reporting and a stronger judge-validation check before trusting the exact numbers.","tokens_in":15995,"tokens_out":2846,"would_cite":true,"duration_ms":26186,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Prompting LLMs with scientific-sounding summaries of (even fabricated) studies makes them produce markedly more biased and toxic responses.","keywords":["LLM jailbreak","persuasion attack","stereotypical bias","scientific language","paper fabrication","toxicity","authority principle","multi-turn dialogue"],"falsifier":"Score the target models' outputs with human raters who are blind to whether the generating prompt contained scientific citations, and compare their scores with GPT-4's. If blind human scores show no increase under the scientific condition, the reported vulnerability is largely an artifact of the judge rather than a property of the target models.","tokens_in":14978,"feed_emoji":"🧪","tokens_out":6630,"duration_ms":53506,"temperature":0.7,"pith_summary":"This paper tries to establish that a scientific veneer is itself a jailbreak vector for current large language models. The authors build prompts that summarize real psychology and social-science papers as evidence that stereotypes are beneficial, then ask models to produce follow-up responses with strong, harmful bias. Across GPT-4o, GPT-4, Llama 3.1, Gemini, and Cohere models, bias and toxicity scores rise substantially compared with zero-shot, role-play, or DAN baselines. They also show the attack needs no scholarly knowledge: an LLM can fabricate fake paper titles and abstracts, and those fabricated summaries persuade target models nearly as well. If true, scientific-sounding authority, not logical argument, is enough to override safety training.","feed_headline":"Fake academic papers make top LLMs more biased","feed_subtitle":"Summaries of invented studies endorsing stereotypes raise bias and toxicity across GPT-4o, Llama, Gemini, and Cohere.","key_machinery":"The mechanism is a two-stage persuasion pipeline built on the authority principle from persuasion research. Stage one produces a persuasive scientific summary: real papers, or LLM-fabricated research ideas with added venue and author metadata, are summarized by GPT-4o into a multi-document summary of the benefits of stereotypical bias. Stage two places that summary in the system message and instructs the target model to write a follow-up response containing strong, harmful bias, including a chain-of-thought rationale explaining why the bias is beneficial. The summary is the load-bearing object: it converts an otherwise-refused request into an authoritative scientific context that the target model complies with.","core_discovery":"In the authors' telling, the central discovery is that authority cues carried by scientific language can be weaponized: real studies on stereotype accuracy, cognitive heuristics, and stereotype boost are deliberately misread as endorsements of harmful bias, summarized by an LLM, and placed in the system message. Target models then generate biased continuations in a multi-turn dialogue, with average bias scores on a 0-4 scale rising from 0 under baselines to values around 1.1-3.3 depending on model and dataset, and toxicity also increasing. When the summary is instead built from fabricated papers invented by GPT-4o, with plausible author and venue metadata added, the effect persists and sometimes grows (for example, GPT-4o's bias score rises from 1.71 to 2.56 on StereoSet). The paper further reports that omitting author names or venues lowers persuasion for some models, that bias compounds over dialogue turns, and that two mutation-based defenses, rephrasing and retokenizing, fail to neutralize the attack.","pith_inferences":["Beyond the paper's claims, the same pipeline may transfer to other authoritative-sounding text domains, such as legal, medical, or financial documents; testing those corpora would show whether the vulnerability is specifically about science or about authority cues in general.","Because the key measurement uses GPT-4 as a judge of bias, a blind human-rater evaluation could separate whether the scientific framing changes the target model's behavior or merely shifts how the judge scores that behavior.","If scientific corpora reinforce this vulnerability, safety training that teaches models to describe literature without endorsing its conclusions might generalize better than prompt filters; that is a testable mitigation direction.","The fabricated-paper result suggests citation-verification tools, which check whether cited studies actually exist, could serve as a concrete defense even though the paper only proposes fact-checking in general terms."],"forward_implications":["Scientific-sounding context alone is enough to bypass safety training that resists direct requests, role-play, and DAN-style jailbreaks in several current models.","The attack can be automated end-to-end: a single LLM can invent the fake papers, summaries, and metadata, so no specialist knowledge or curated papers are required.","Adding credibility metadata such as author names and venues strengthens the effect for some models, so removing that metadata is not a reliable defense.","Bias escalates across dialogue turns, meaning multi-turn conversations do not currently re-anchor the model to neutral behavior.","Rephrasing and retokenization defenses are ineffective and can even raise bias scores, so standard prompt-mutation defenses do not blunt this vector."],"supporting_citations":[{"why":"Supplies the authority principle that motivates using scientific language as persuasion.","marker":"Cialdini, 2007"},{"why":"Provides the StereoSet dataset of neutral context sentences with gender, race, religion, and profession bias labels used for evaluation.","marker":"Nadeem et al., 2020"},{"why":"Shows that LLM-as-a-judge bias scores align with human judgments, forming the basis for the paper's main evaluation metric.","marker":"Kumar et al., 2024"},{"why":"Motivates the fabricated-paper study by showing that literature-based discovery can generate new scientific ideas.","marker":"Swanson, 1986"},{"why":"Inspires the chain-of-thought rationale request that makes target models explain why their biased response is beneficial.","marker":"Wei et al., 2022"},{"why":"One of the real papers deliberately reinterpreted as evidence that heuristics and biases are beneficial.","marker":"Tversky and Kahneman, 1974"},{"why":"One of the real papers on stereotype performance boosts that is summarized as endorsing harmful stereotypes.","marker":"Shih et al., 2002"},{"why":"Prior persuasion-based jailbreak work whose approach this paper extends and compares against.","marker":"Zeng et al., 2024"},{"why":"Provides the mutation-based defenses, rephrasing and retokenization, that the paper tests and finds ineffective.","marker":"Jain et al., 2023"}],"fun_headline_variants":["Scientific jargon hides jailbreaks that bias LLMs","Studies misread as endorsements turn LLMs toxic","Malicious prompts sneak in via fake science","LLMs swayed by fabricated scientific studies","Fake research boosts LLM bias and toxicity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central evidence assumes that GPT-4's bias-score judgment is impartial; if the scientific framing of the prompt also sways the judge, the reported rise in bias could be partly a measurement artifact.","fun_headline_variants_meta":{"raw":{"variants":["Scientific jargon hides jailbreaks that bias LLMs","Studies misread as endorsements turn LLMs toxic","Malicious prompts sneak in via fake science","LLMs swayed by fabricated scientific studies","Fake research boosts LLM bias and toxicity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000439,"raw_usage":{"total_tokens":2242,"prompt_tokens":973,"completion_tokens":1269,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":1199}},"tokens_in":589,"tokens_out":1269,"duration_ms":8253,"temperature":1.0,"reasoning_tokens":1199,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:23:30.720638+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Score the target models' outputs with human raters who are blind to whether the generating prompt contained scientific citations, and compare their scores with GPT-4's. If blind human scores show no increase under the scientific condition, the reported vulnerability is largely an artifact of the judge rather than a property of the target models.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the authority principle that motivates using scientific language as persuasion."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the fabricated-paper study by showing that literature-based discovery can generate new scientific ideas."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"One of the real papers on stereotype performance boosts that is summarized as endorsing harmful stereotypes."}],"review_version":1}