{"id":"0195936c-4180-4df8-9b02-19747d58c4d4","arxiv_id":"2510.03905","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"Yearly word embeddings show Japanese gender stereotypes for Home, Work, and Politics all grew more female-associated from 1900 to 1999, but the claimed 1945 reversals are not formally tested and may reflect corpus-wide frequency shifts rather than attitude change.","lead":"This paper trains yearly word-embedding models on a century of Japanese books and magazines (1900–1999) and measures how strongly Home, Work, Politics, and 18 occupations associate with female versus male words. It reports that all three domains became more female-associated, but the headline claim of a sharp post-1945 reversal is contradicted by the abstract’s own statement about Home and is complicated by a corpus-wide frequency trend.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim is contradicted by the paper's own results: Home is reported as stable in the abstract but shows a significant positive trend (Table 3), and the '1945 reversal' has no breakpoint test.","rationale":"The reader's rejection is well supported, and my stress-test identifies the same central-claim failure but locates the most load-bearing problem slightly differently. The reader's weakest_assumption focuses on the frequency confound documented in Appendix A; that is a serious threat to any version of the conclusion. However, the more immediate and decisive problem is internal inconsistency: the abstract's claim that Home is stable is contradicted by the paper's own Table 3 and Section 5.1.1, and no breakpoint test supports the asserted 1945 reversal. There is no need to invoke external confounds to show the central claim as stated is unsupported. The concrete test—a piecewise regression with a 1945 breakpoint on both raw and adjusted series—would settle whether any defensible version of the 'reversal' survives. Given the manuscript's own evidence contradicts its headline, the correct verdict remains REJECT; my read does not call for a change to the reader's verdict. I mark agreement as 'partial' because the reader and I converge on rejection but for somewhat different primary reasons: internal contradiction versus frequency confound.","tokens_in":20227,"tokens_out":3234,"duration_ms":29069,"concrete_test":"Fit a piecewise linear model with a breakpoint at 1945 to the yearly Home, Work, and Politics stereotype series (using the bootstrap distributions for uncertainty), and test (i) whether the post-1945 slope minus pre-1945 slope differs from zero for each domain, and (ii) whether Home's post-1945 slope is significantly different from zero. If Home shows a significant positive post-1945 slope or no significant slope change relative to prewar, the abstract's 'stable Home/no significant changes' claim is directly rejected. Run the same test on the adjusted scores (S − corpus average) and on a matched control set controlling for gender-word frequency; if the reversal disappears under frequency control, the 1945-specific conclusion is an artifact of corpus statistics rather than semantic association.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, as stated in the supplied abstract, is that Home-domain female stereotype values are stable with no statistically significant changes, while Work and Politics exhibit significant trend reversals around 1945. The full text contradicts this: Section 5.1.1 states that 'words across all three domains show a rise in female stereotypes beginning around 1945,' and Table 3 reports positive trend coefficients for all three domains, all p<0.01, with Home showing the largest per-year increase (0.000128). Even the full-text abstract says 'Home also became more female-stereotyped.' No formal breakpoint test is reported for a 1945 reversal; the only trend evidence is whole-century linear regressions. Moreover, Section 5.3 shows a corpus-wide upward trend (Table 5: 0.000137) steeper than Work and Politics, and the adjusted analysis in Table 6 shows only Work has a significant negative relative trend (p<0.05), meaning Work is increasing more slowly than the corpus average—not reversing in any absolute sense. Thus the differential 'stable Home vs. changing public domains' conclusion is not derivable from the reported evidence and appears to rest on an earlier version of the analysis. Appendix A compounds this by documenting r=0.60 between the corpus-wide stereotype measure and the relative frequency gap of gender words, driven by 彼(he), with the authors conceding this 'may have influenced female stereotype values calculated through the embedding models.' The central claim fails on internal consistency before external-validity questions are even considered.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains year-specific SGNS word embeddings on the Japanese NDL n-gram corpus (1900–1999) and computes a WEAT-style female-stereotype score for three domains (Home, Work, Politics) and 18 occupations. It reports that all three domains show significant positive whole-century trends in female stereotype, that the Home domain remains the most female-stereotyped throughout, and that occupational stereotype values correlate moderately with census gender participation rates. The manuscript also reports a corpus-wide upward trend in the same stereotype measure and an adjusted analysis that removes this corpus average.","tokens_in":20541,"tokens_out":5265,"duration_ms":46002,"significance":"If the results are interpreted carefully, the study offers a useful longitudinal, non-Western benchmark for embedding-based stereotype measurement, and the public release of the trained yearly embeddings is a concrete asset. The occupation–census validation (overall r = 0.413, with significant correlations for doctors and dentists) is an important external check. However, the headline claim of a 1945-specific reversal with a stable Home domain is not supported by the statistics actually reported; the body's own findings are more modest and internally coherent. The significance of the paper therefore depends on substantial revision of the framing and additional analysis.","major_comments":[{"comment":"The abstract states that 'In the Home domain, female stereotype values remain stable over time, showing no statistically significant changes.' This is directly contradicted by Section 5.1.1 and Table 3, which report a Home-domain trend coefficient of 0.000128 (SE 1.42e-5, p < 0.01). The full-text abstract itself says 'Home also became more female-stereotyped.' The central domain-specific contrast of the abstract is thus not supported by the paper's own results. This is not a wording issue: it reverses the main claimed finding.","section":"Abstract"},{"comment":"The paper claims 'significant trend reversals in female stereotype values around 1945, shifting from negative prewar trends to positive postwar trends' for Work and Politics. No statistical breakpoint test is provided. Tables 3 and 5 report only whole-century linear regressions; Figure 2 shows a visual inflection but the paper does not present a Chow test, segmented regression, or pre/post 1945 slope comparison. Moreover, the adjusted analysis in Table 6 shows that only Work has a significant (p < 0.05) negative relative trend; Politics does not. The reported evidence supports a gradual increase, not a statistically tested reversal at 1945.","section":"§5.1.1, Table 3, Table 6"},{"comment":"Appendix A documents r = 0.60 between the corpus-wide stereotype measure and the relative frequency gap between female and male gender words, and attributes the pattern largely to the rise and fall of 彼 (he). The authors concede this 'may have influenced female stereotype values calculated through the embedding models.' Because the same cosine-difference measure is applied to all target domains, this frequency-driven change in the attribute words' embeddings could affect the domain trends, including the post-1945 rise. Subtracting the corpus-wide average in §5.3 does not remove a non-uniform effect of the frequency gap across domains. The central semantic interpretation therefore remains at risk.","section":"Appendix A, §5.3"},{"comment":"The embedding-quality checks show that yearly similarity and association scores correlate strongly with vocabulary size (r = 0.77 and 0.82). Since the paper compares stereotype values across years, systematic variation in model quality over time—especially around the periods of corpus-size collapse—could create or obscure trends. This is not mentioned in the Limitations section, and the paper does not test whether the main results are robust to year-specific model quality.","section":"Appendix D"}],"minor_comments":[{"comment":"There are many typos (e.g., 'centry', 'streotype', 'femle', 'vocaburarly', 'eacy year's word embeddding models'). A thorough proofread is needed.","section":"Throughout"},{"comment":"The sentence 'We used SGNS of its widespread adoption' should read 'We used SGNS because of its widespread adoption' or similar.","section":"§4.2"},{"comment":"Figures 2, 6, and B3 repeat 'over the centry' and 'The red vertical line corresponds to 1945.' Also, Figures 5 and 6 claim confidence intervals are shown but 'not visible'; consider plotting them with a meaningful scale or omitting the shaded region.","section":"Figure captions"},{"comment":"The heading 'Corpus-wide trend of increasing female stereotype?' is a question rather than a finding. Rephrase as a statement or a clearly labeled exploratory discussion.","section":"§5.3.1"},{"comment":"Table 6 uses a dagger (†) for 0.01 ≤ p < 0.05, but the note in the table says only '*' is used. The note should list both markers.","section":"§4.2 / Table 3 note"}],"recommendation":"major_revision","confidential_remarks":"The paper appears to have two different abstracts: the arXiv/provided abstract claims the Home domain is stable, while the full text abstract and body conclude Home also became more female-stereotyped. This version mismatch should be reconciled as a matter of editorial integrity. The body's actual findings, if reframed, are publishable; but the reversal claim requires formal breakpoint testing and the frequency confound needs a substantive response, not just a correlational appendix note."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The first thing you should know: this is a serious empirical study with a new result, but the paper in its current form does not support the claim in its official abstract. The body shows all three domains—Home, Work, Politics—rising significantly over the century, with Home rising the fastest. The official abstract says Home is stable and reserved the reversal for Work and Politics. That is not a minor wording issue; it is the difference between the headline finding and what the data actually show. The full-text abstract even says \"Home also became more female-stereotyped,\" so the contradiction is internal.\n\nWhat the paper does well: it is the first embedding-based account of gendered discourse across Japan's prewar–postwar transition, built on the NDL Ngram corpus with diachronic embeddings trained per year. The authors use a historical morphological dictionary, release the models, and attempt external validation against census data on occupational gender composition. The pooled occupation–census correlation (r=0.413) is a meaningful positive check. The adjusted analysis in Section 5.3 is also a good instinct, even if the interpretation goes sideways.\n\nThe soft spots are consequential. First, there is no formal breakpoint test for the \"around 1945\" reversal; the claim rests on visual inspection of Figure 2. Second, the adjusted analysis shows only Work has a significant negative relative trend, meaning Work is rising more slowly than the corpus average—not reversing in any absolute sense. That directly contradicts the \"Work reversed, Home stable\" narrative. Third, and most damaging, Appendix A documents a correlation of r=0.60 between the corpus-wide stereotype measure and the relative frequency gap of gender words, driven by 彼 (he). The authors concede this \"may have influenced female stereotype values.\" Since the same gender word lists are used for the domain measures, this confound is load-bearing, not a side note. The corpus-wide trend is essentially indistinguishable from a pronoun-frequency artifact,\nand the domain trajectories may share that problem.\n\nDespite these flaws, the paper is not a throwaway. The data work is careful, the limitations section is honest, and the underlying pipeline could be salvageable with a corrected abstract, formal breakpoint tests, and frequency-controlled robustness checks. I would send it to peer review, not desk reject it, because the topic matters and the empirical contribution is genuinely new. But I would expect major revision before publication, and I would not cite the current version in my own work.\n\nFor a reading group, it is useful as a case study in how historical embedding results can be confounded by corpus statistics. The citation pattern looks fair; the authors engage with the relevant Garg, Kozlowski, Jones, and Charlesworth literature. This is a paper where a senior referee could make a real difference by pushing on the measurement assumptions.","headline":"The empirical core is real and worth knowing about, but the abstract contradicts the paper's own results on the Home domain, the 1945 reversal is asserted without a test, and the frequency confound in Appendix A cuts against the central measurement.","tokens_in":21063,"tokens_out":2794,"would_cite":false,"duration_ms":27741,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Female stereotypes rise across Japanese domains after 1945","keywords":["word embeddings","gender stereotypes","Japan","historical corpora","linguistic change","gender roles","census validation","postwar reforms"],"falsifier":"Rerun the stereotype calculation with 彼(he) and 彼女(she) removed from the male and female attribute lists, or with the attribute lists frequency-matched year by year. If the post-1945 rise in Home, Work, and Politics disappears or reverses, then the paper's finding is an artifact of lexical-frequency shifts in Japanese writing rather than a genuine change in gender stereotypes.","tokens_in":20101,"feed_emoji":"📈","tokens_out":9836,"duration_ms":71082,"temperature":0.7,"pith_summary":"The paper asks whether Japan's post-1945 legal reforms—granting women the vote, constitutional equality, and labor protections—actually moved gendered language. It trains separate word-embedding models for each year from 1900 to 1999 on a large corpus of digitized Japanese books and magazines, then measures how close domain words sit to female versus male terms. The measurements show all three domains—Home, Work, Politics—becoming more female-associated after 1945; Work and Politics flip from male- to female-stereotyped by the 1970s, while Home remains the most female-stereotyped domain throughout. A census validation shows occupations with more women have higher female stereotype scores, so the language measure tracks real demographics. The paper interprets this as women's public roles being added to, not replacing, domestic ones.","feed_headline":"Female stereotypes rise across Japanese domains after 1945","feed_subtitle":"Word-embedding analysis of 100 years of books ties the shift to postwar reforms and to women's real share of occupations.","key_machinery":"The central mechanism is a series of yearly skip-gram word-embedding models (vector representations learned from word co-occurrence in a five-word context window) trained on Japanese n-gram frequency data. The paper computes a female stereotype value for each target word as the mean cosine similarity to female attribute words minus the mean cosine similarity to male attribute words, averaged over the words in a domain or occupation list, with bootstrap confidence intervals. To separate genuine domain change from a decade-long rise in female association across the entire vocabulary, it subtracts the corpus-wide average stereotype each year to produce an adjusted score. This adjusted-score mac","core_discovery":"The paper's central claim is that gendered cultural discourse in Japan shifted unevenly across the prewar-postwar transition, and that this shift is measurable in word co-occurrence statistics. Using yearly word embeddings, the authors show that female stereotype values—the mean cosine distance of a domain's words to female versus male attribute words—rose in the Home, Work, and Politics domains after 1945. Work and Politics crossed from male-stereotyped to female-stereotyped territory around 1970, while Home remained the most female-stereotyped domain for the entire century. The paper also reports a statistically significant positive correlation between the female stereotype score of an occ","pith_inferences":["The abstract states that the Home domain 'remains stable over time, showing no statistically significant changes,' but the full text's Table 3 reports a significant positive trend for Home; readers should treat the abstract as an unreliable summary of the paper's own results.","If the frequency of the pronoun 彼(he) drove the corpus-wide trend, the post-1945 reversal in Work and Politics could partly reflect a literary-style shift (adoption then abandonment of Western-style third-person pronouns) rather than a change in gender beliefs; removing 彼 and 彼女 from attribute lists would test this.","The same adjusted-score approach could be applied to other digitized historical corpora (e.g., newspapers, political speeches) to see whether the 1945 inflection is specific to books or appears across registers.","The occupation-level census correlations are thin (at most nine data points per occupation); the stronger pooled correlation suggests future work should prioritize occupations with richer time series."],"forward_implications":["If the measurements hold, postwar legal changes left a measurable trace in everyday Japanese prose: Work and Politics terms became more female-associated within roughly two decades, and the Home domain never lost its female association.","The persistent, even increasing, female stereotype in the Home domain suggests the cultural model was additive (homemaker plus worker/politician) rather than substitutive.","Because the corpus-wide average also rises, raw stereotype scores overstate domain-specific change; the adjusted score should be the default in future work on this corpus.","The occupation-census correlation indicates that when demographic data are sparse, embedding-based stereotype measures can serve as a proxy for women's actual presence in occupations.","The rise in the corpus-wide score beginning around 1966 implies that a global linguistic drift toward female association exists in Japanese, separate from any particular domain."],"fun_headline_variants":["Japan's post-1945 gender shift: work and politics lead","Word embeddings show Japan's gender shift post-1945","Japan's gendered work and politics reversed after 1945","Japan's book gender stereotypes shift unevenly around 1945","Post-1945 Japan: gendered books shift in work and politics"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central measurement assumes that cosine-distance differences to female versus male gender words reflect semantic gender association, not shifts in how frequently those gender words occur—especially the changing frequency of the pronoun 彼(he), which correlates with the corpus-wide stereotype trend at r=0.60.","fun_headline_variants_meta":{"raw":{"variants":["Japan's post-1945 gender shift: work and politics lead","Word embeddings show Japan's gender shift post-1945","Japan's gendered work and politics reversed after 1945","Japan's book gender stereotypes shift unevenly around 1945","Post-1945 Japan: gendered books shift in work and politics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000931,"raw_usage":{"total_tokens":3853,"prompt_tokens":806,"completion_tokens":3047,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":2961}},"tokens_in":550,"tokens_out":3047,"duration_ms":69346,"temperature":1.0,"reasoning_tokens":2961,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T11:33:42.624737+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the stereotype calculation with 彼(he) and 彼女(she) removed from the male and female attribute lists, or with the attribute lists frequency-matched year by year. If the post-1945 rise in Home, Work, and Politics disappears or reverses, then the paper's finding is an artifact of lexical-frequency shifts in Japanese writing rather than a genuine change in gender stereotypes.","supporting_citations":[],"review_version":1}