{"id":"adc2cd39-d7c5-4366-9cdc-4685a6f74a62","arxiv_id":"2501.01987","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Sora-generated videos associate stereotyped occupations, behaviors, and appearances with specific genders, based on frequency counts from 120 clips.","lead":"This paper checks whether OpenAI's Sora video generator reproduces gender stereotypes by generating 120 short videos from 12 prompts and counting the apparent gender of each character. The counts show strong stereotypical associations, for example nurses and secretaries as female and CEOs and muscular people as male, but the evidence is based on 10 videos per prompt with no statistical testing.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The counts are unverifiable and underpowered: no Sora access details, no released videos or annotations, no inter-rater reliability, and n=10 per prompt cannot support 'exclusively' or 'significant'.","rationale":"I read the paper as a small, transparent audit: generate videos from neutral prompts, count apparent gender, report stereotypes, and test prompt-level debiasing. The strongest part is the direct-debiasing demonstration; that is internally consistent and useful. The central claim, however, is not the debiasing result; it is the claim that Sora's default outputs are gender-stereotyped to a significant degree. That claim depends on (a) the videos actually being Sora outputs, (b) the gender counts being accurate and reproducible, and (c) the sample being large enough to support the qualitative words used. All three are currently uncheckable. The paper itself acknowledges in Section 5 that the number of runs is limited and generalizability is constrained, but it does not supply the data or error bars needed to interpret the counts. I do not think this requires rejection; the study could be made acceptable by releasing artifacts and adding statistics, exactly the conditions the reader proposed. I therefore leave the verdict unchanged, while noting that the conditionality is not optional.","tokens_in":5019,"tokens_out":2904,"duration_ms":30506,"concrete_test":"Obtain from the authors (1) the exact Sora access method, model version, sampling configuration, and timestamps; (2) the 120 raw videos or representative frames; and (3) independent gender annotations by at least two annotators using a written protocol. Then recompute per-term gender counts, Cohen's kappa, and exact binomial 95% confidence intervals. The central claim is secure only if kappa >= 0.8, the CI for claimed 'exclusive' categories does not overlap a 50/50 split, and at least one auditor can reproduce the reported counts from the released data. If any of these fail, the 'significant' conclusion should be downgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim — that Sora disproportionately associates genders with stereotypical terms — rests entirely on frequency counts from 120 videos, but the paper provides no way to verify those counts. Section 2 says only that each prompt was run 10 times and the videos were 'analyzed to identify the gender'; it does not state how Sora was accessed, which model version or sampling parameters were used, when generation occurred, or whether outputs were deterministic. More importantly, the videos and raw annotations are not shared, and no annotation protocol, annotator count, or agreement measure is given. This makes the headline result non-reproducible. A second, compounding issue is statistical: with 10 generations per term, 'exclusively male' (e.g., Muscular) has a 95% confidence interval of roughly 69–100% for the male proportion; true female rates as high as ~26% are not excluded by the data. 'Ugly' being 'balanced' is claimed with no test. While the direct-debiasing examples (Figure 3) are plausible illustrative evidence, the quantitative claim of 'significant' bias is not supported by the reported evidence. Because the conclusion is entirely downstream of these unverified, underpowered counts, this is the load-bearing weak point.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a case study of gender bias in OpenAI's Sora text-to-video model. The authors generated 120 short videos from 12 gender-neutral prompts spanning Appearance, Behavior, and Occupation categories (10 runs per prompt), then manually judged the perceived gender of the primary character in each video. They report frequency counts indicating that Sora associates 'Attractive', 'Frail', 'Shy', 'Emotional', 'Nurse', and 'Secretary' predominantly or exclusively with women, and 'Muscular', 'Confident', 'Rational', and 'CEO' predominantly or exclusively with men, while 'Ugly' and 'Doctor' appear balanced. They also test two prompt-level debiasing strategies: direct gender specification (which works) and indirect unbiased phrasing (which reportedly fails, with 10/10 male CEOs for 'A CEO working. It could be male or female'). The paper concludes that Sora exhibits significant gender bias reflecting societal stereotypes, and recommends future work on multi-model, multi-dimensional bias analysis.","tokens_in":5232,"tokens_out":2648,"duration_ms":28587,"significance":"If the reported counts are accurate, this is one of the first published empirical investigations of gender bias in a modern text-to-video generation model, and it would usefully extend prior image-generation bias findings (e.g., DALL-Eval, Hamidieh et al.) to the video domain. The choice of simple, neutral prompts and the inclusion of a debiasing experiment are sensible first steps, and the authors explicitly acknowledge several limitations. However, the paper ships no code, data, or annotation artifacts, and its central quantitative claims are not underpinned by statistical inference or inter-rater reliability measures. The value of this manuscript is therefore preliminary and qualitative: it provides a plausible demonstration that Sora's default outputs align with gender stereotypes, but it does not currently support the strong quantitative language used in the abstract and conclusions.","major_comments":[{"comment":"The central empirical claim rests on frequency counts from 120 videos, but the manuscript does not state how Sora was accessed, which model version or sampling parameters were used, when the generations occurred, or whether any randomness/seed control was applied. Without this information and without releasing the generated videos or raw annotations, the reported counts are not reproducible. The authors should provide a detailed generation protocol and make the video corpus and annotation data available, or clearly frame the results as an illustrative case study rather than a reproducible measurement.","section":"Section 2, Video Generation"},{"comment":"The gender judgments are made by the researchers, but no annotation protocol, number of annotators, or inter-rater reliability (e.g., Cohen's kappa) is reported. Since 'gender' in generated videos is a visual interpretation that can be ambiguous, the reported male/female frequencies could reflect annotator expectations as much as model output. The authors should describe the annotation instructions, use multiple independent annotators, and report agreement statistics.","section":"Section 2, Video Generation and Figure 1"},{"comment":"With only 10 runs per prompt, categorical claims such as 'Muscular was exclusively associated with males' and 'Secretary was entirely associated with females' are not statistically justified. For a 10/10 count, the 95% confidence interval for the true proportion extends from approximately 69% to 100%, so the data are consistent with female occurrences up to about 30% that simply did not appear in the sample. The authors should report confidence intervals or exact binomial tests, and replace 'exclusively', 'entirely', and 'overwhelmingly' with language calibrated to the sample size.","section":"Section 3, Results and Figure 1"},{"comment":"The indirect debiasing result—'A CEO working. It could be male or female' produced male CEOs in all 10 runs—is used to conclude that 'indirect strategies alone cannot overcome stereotypes in the generation process.' With n=10, this observation is weak: under a model that generates male CEOs with probability 0.85, the chance of observing 10 males is about 20%. A statistical test, a larger number of runs, or multiple indirect-prompt variants is needed before drawing a strong conclusion about the failure of indirect debiasing.","section":"Section 4, Debiasing"},{"comment":"The authors acknowledge in Section 5 that 'the experiments are based on a limited number of runs, which could impact the reliability of the results due to potential variability in the model's outputs' and that increasing runs 'would enhance the statistical significance.' This acknowledged limitation directly contradicts the abstract's claim of 'significant evidence of bias' and 'disproportionately associates.' The paper should either supply the missing statistical analysis, or revise the abstract and conclusion to describe the findings as suggestive and preliminary.","section":"Section 5, Limitations and Abstract"}],"minor_comments":[{"comment":"The screenshots in Figure 2 are not individually labeled with the corresponding prompt or the judged gender, making it difficult for the reader to map each image to the reported counts; consider adding subcaptions such as '(e) Confident — male depiction'.","section":"Figure 2"},{"comment":"The authors state that prompts were 'designed to be simple, direct, and neutral' but do not provide the full list of exact prompts (e.g., whether 'A doctor working' vs. 'An attractive person' were uniformly phrased). Listing all twelve prompts verbatim would improve replicability.","section":"Section 2, Prompt Formation"},{"comment":"Reference [8] is a review paper on Sora rather than the official OpenAI technical report; citing the primary source (OpenAI, 'Video generation models as world simulators', 2024) would be more accurate for the architectural claims in Section 2.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a straightforward case study with useful illustrative material, but its central quantitative claims are not currently supported by the evidence provided. The lack of reproducible data and statistical tests is a standards issue for a journal, not just a cosmetic concern. I would not reject the paper outright, because the qualitative direction is plausible and consistent with prior text-to-image bias literature, and the authors could fix the load-bearing weaknesses by adding statistical analysis, annotation reliability, and either a data-release statement or a clearly exploratory framing. The journal should weigh whether a single-model, 12-prompt, n=10-per-prompt study without released artifacts meets its bar for empirical contributions; if not, the authors should be encouraged to either strengthen the evidence or resubmit to a workshop-oriented venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a small, transparent audit of Sora that finds expected stereotyped defaults, plus a genuinely interesting negative result on indirect debiasing. It deserves a serious referee, but the headline counts need to be verifiable before I'd trust them.\n\nWhat's new: the twelve-prompt frequency counts for Sora, and the observation that explicit gender prompts override defaults while the soft prompt \"A CEO working. It could be male or female\" produced male CEOs in 10/10 runs. That failure of indirect debiasing is the most valuable thing in the paper. The audit method itself is a routine adaptation of Hamidieh et al., no new technique, but that's acceptable for a case study.\n\nThe qualitative direction matches earlier text-to-image bias results, so nothing here is surprising. The screenshots and discussion of layered stereotypes (shy people young, secretaries on landline phones) are suggestive illustrations, not measured findings, and the paper mostly treats them that way.\n\nNow the soft spot, and it's load-bearing: the claim of \"significant evidence of bias\" is not supported by the reported data. Ten runs per prompt, no confidence intervals, no significance tests, no annotation protocol, no inter-rater reliability, and no released videos or raw annotations. With n=10, \"exclusively male\" is compatible with a true female rate up to around 26%. The paper's own limitations section acknowledges the small number of runs, so the authors know this, but the abstract still says \"significant\" and the results use words like \"exclusively\" and \"entirely.\" That mismatch matters.\n\nThe stress-test note holds up: the counts are unverifiable. The methodology section says only that videos were \"analyzed\" to identify gender, with no detail on annotators, agreement, Sora access, model version, or sampling parameters. This is fixable, but it has to be fixed before the central claim can be accepted.\n\nOn the citation side, the reference list is fine. The self-citation to the authors' negation-blindness work is relevant and not padding.\n\nRecommendation: send it to peer review, but with conditions. Ask for the videos or raw annotations, a documented annotation protocol with agreement measure, basic statistics (confidence intervals or an exact test), model version and sampling details, and a softening of \"exclusively\" and \"significant\" to match the evidence. The indirect-debiasing result is worth keeping no matter what.","headline":"Small transparent Sora audit with a useful negative debiasing result, but the headline counts lack verification and statistical support.","tokens_in":5779,"tokens_out":1590,"would_cite":false,"duration_ms":17650,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Sora, OpenAI's text-to-video model, renders neutral prompts with stereotyped genders: nurses and secretaries as female, CEOs and muscular people as male.","keywords":["gender bias","text-to-video generation","Sora","stereotypes","AI fairness","vision-language models","prompt-level debiasing","social bias in generative AI"],"falsifier":"Run the same twelve prompts through Sora again and have several independent annotators, blind to the prompts, classify the characters' genders — or apply a published automatic gender classifier to the generated faces. If inter-annotator agreement is low, or if the resulting male/female frequencies diverge from the reported pattern (for example, nurses not predominantly female or CEOs not predominantly male), the paper's central claim would collapse. Because Sora's sampling is stochastic, repeating each prompt many times and reporting the distribution would make this test decisive.","tokens_in":4827,"feed_emoji":"🎬","tokens_out":8831,"duration_ms":81060,"temperature":0.7,"pith_summary":"Text-to-video models are entering real content pipelines, so the demographics they default to matter. This paper tests OpenAI's Sora with twelve gender-neutral prompts — \"A nurse working,\" \"A CEO working,\" \"An attractive person\" — repeated ten times each. Counting the apparent gender of the people in the 120 generated videos, the authors report that Sora reliably maps stereotype-linked terms along traditional gender lines: nurses and secretaries come out as women, CEOs and muscular people as men, shy and emotional people as women, confident and rational people as men. The paper argues that these defaults mirror societal biases in the training data, and that only explicitly naming a gender in the prompt overrides them; an indirect instruction that the person \"could be male or female\" does not.","feed_headline":"Sora renders nurses female and CEOs male from neutral prompts","feed_subtitle":"A 120-video probe finds nurses, secretaries and shy people female; CEOs and muscular people male.","key_machinery":"The mechanism is a minimal probe-and-count setup. Twelve stereotype-linked terms are grouped into three categories — Appearance, Behavior, Occupation — and each is embedded in a simple gender-neutral prompt such as \"A nurse working\" or \"An attractive person.\" Sora generates ten independent five-second videos per prompt; the authors visually classify gender in each video and tabulate male/female frequencies. Those frequency tables are the evidence that Sora's default gender assignment follows stereotypes. A second mechanism tests whether the bias can be steered from the prompt: explicit gender naming (\"A male nurse\") versus an indirect fairness instruction (\"A CEO working. It could be male or female\"), with the two giving opposite results.","core_discovery":"The paper's central claim is that Sora, a state-of-the-art text-to-video model, systematically produces gender-stereotyped content from neutral prompts. In the Appearance category, \"Attractive\" and \"Frail\" are predominantly rendered as female while \"Muscular\" is exclusively male and \"Ugly\" balanced. In the Behavior category, \"Confident\" and \"Rational\" skew male, \"Shy\" female, and \"Emotional\" overwhelmingly female. In the Occupation category, \"Nurse\" and \"Secretary\" are generated exclusively as female, \"CEO\" overwhelmingly male, and \"Doctor\" roughly balanced. The paper also notes layered secondary stereotypes in the videos — older people shown as frail, shy people as young, confident people in formal work attire, secretaries on landline phones — and reports that direct prompt-level gender specification succeeds where indirect neutral instruction fails. The conclusion is that Sora inherits and amplifies stereotypes from its training data, so text-to-video models need their own bias audits.","pith_inferences":["Editorial extension: the reported frequencies describe one stochastic sampling trace of a closed model, so production deployments could drift from this distribution unless the generation settings are fixed and monitored.","Editorial extension: a natural next experiment is to hold these twelve prompts fixed and run them through several text-to-video models, comparing male/female distributions; that would show whether the pattern is specific to Sora or common to web-scale video generators.","Editorial extension: the analysis is binary only, so a fuller bias audit would add skin-tone, age, and non-binary gender presentation to the same prompts; the paper's own \"layered stereotypes\" observations suggest these dimensions interact with gender.","Editorial extension: the indirect-debiasing failure on \"CEO\" invites a systematic search over blind fairness phrasings, plural subjects, and balanced person descriptions to find whether any prompt-only intervention can move Sora's default."],"forward_implications":["Neutral prompts to Sora do not produce neutral demographics: downstream creative, educational, or journalistic use inherits a gendered frame (nurse as woman, CEO as man).","Explicitly naming a gender in the prompt reliably overrides the default, giving content creators a concrete mitigation that works at generation time.","A one-line fairness instruction that does not name gender is ineffective, at least for this model and prompt set, so prompt-level debiasing has a hard limit.","Because Sora is trained on web-scale data, similar stereotype patterns should be expected in other text-to-video models and should be audited before deployment.","Fairness fixes for text-to-video will need more than prompt rewriting, such as fairness-aware training or post-hoc content controls."],"supporting_citations":[{"why":"Describes Sora as a diffusion-transformer text-to-video model; it fixes the black-box system whose outputs are probed.","marker":"[8]"},{"why":"Supplies the Appearance/Behavior/Occupation prompt categories and the idea of probing implicit social biases in vision-language models.","marker":"[13]"},{"why":"Shows DALL-E and Stable Diffusion exhibiting gender and skin-tone biases, the pattern this paper looks for in video.","marker":"[12]"},{"why":"Demonstrates stereotypical gender, profession, and race bias in language models, motivating the same test in text-to-video.","marker":"[11]"},{"why":"Surveys social bias in vision-language models and supports the claim that biases come from internet-scale training data.","marker":"[15]"},{"why":"Documents a persistent stereotype association in GPT-3, an example of training-data prejudice carried into generative output.","marker":"[9]"}],"fun_headline_variants":["Sora turns neutral prompts into gender-stereotyped videos","Study: Sora renders nurses female, CEOs male from neutral text","Text-to-video model Sora amplifies gender stereotypes in outputs","Neutral prompts in Sora yield gendered job roles and traits","Sora's videos show gender bias from neutral prompts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the authors' visual judgments of each video character's gender are accurate and consistent; the paper reports no annotation protocol, no second annotator, no agreement measure, and no released video data.","fun_headline_variants_meta":{"raw":{"variants":["Sora turns neutral prompts into gender-stereotyped videos","Study: Sora renders nurses female, CEOs male from neutral text","Text-to-video model Sora amplifies gender stereotypes in outputs","Neutral prompts in Sora yield gendered job roles and traits","Sora's videos show gender bias from neutral prompts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000742,"raw_usage":{"total_tokens":3259,"prompt_tokens":840,"completion_tokens":2419,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":456,"completion_tokens_details":{"reasoning_tokens":2333}},"tokens_in":456,"tokens_out":2419,"duration_ms":17455,"temperature":1.0,"reasoning_tokens":2333,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:01:00.039077+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same twelve prompts through Sora again and have several independent annotators, blind to the prompts, classify the characters' genders — or apply a published automatic gender classifier to the generated faces. If inter-annotator agreement is low, or if the resulting male/female frequencies diverge from the reported pattern (for example, nurses not predominantly female or CEOs not predominantly male), the paper's central claim would collapse. Because Sora's sampling is stochastic, repeating each prompt many times and reporting the distribution would make this test decisive.","supporting_citations":[{"cited_title":"A survey on video diffusion models","cited_arxiv_id":null,"evidence_quote":"Supplies the Appearance/Behavior/Occupation prompt categories and the idea of probing implicit social biases in vision-language models."},{"cited_title":"Generatect: Text -conditional generation of 3d chest ct volumes","cited_arxiv_id":null,"evidence_quote":"Shows DALL-E and Stable Diffusion exhibiting gender and skin-tone biases, the pattern this paper looks for in video."},{"cited_title":"The Current State of Artificial Intelligence and Its Intersection With Radiology","cited_arxiv_id":null,"evidence_quote":"Demonstrates stereotypical gender, profession, and race bias in language models, motivating the same test in text-to-video."},{"cited_title":"Text2video -zero: Text-to-image diffusion models are zero-shot video generators","cited_arxiv_id":null,"evidence_quote":"Documents a persistent stereotype association in GPT-3, an example of training-data prejudice carried into generative output."}],"review_version":1}