{"id":"064df3ee-9140-430d-ad2a-16b6dc06f246","arxiv_id":"1908.01352","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An automated analysis of brand ads on social media finds White models overrepresented, Black models underrepresented, a preference for same-race cross-sex pairings, and higher smiling scores for female models.","lead":"This paper analyzes 85,957 advertising images from 73 top brands on Instagram and Facebook to measure how often different genders and races appear, pair up, and smile. It finds White models dominate, Black models are underrepresented, same-race male-female pairings occur more often than chance, and female models smile more than male models.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated Face++ labels on this specific corpus support every headline number; the paper's own Gender Shades citation makes the missing stratified validation the pivotal risk.","rationale":"The reader's weakest assumption is Face++ accuracy; I agree that this is the gate. I do not see a fatal internal flaw in the permutation null model; it is a reasonable baseline, and the paper is careful to write 'inferred as' for every demographic label. The issue is that the paper's own cited evidence (Gender Shades) and its own confidence threshold create a specific, unquantified selection and measurement risk. The central claim would be true if, say, Face++ error were non-differential by race and gender; but the cited literature suggests differential error, so the burden is on validation. I keep the verdict at CONDITIONAL rather than REJECT because the observed White majority is large enough that moderate classifier error is unlikely to erase it; the proposed check would settle whether the finer claims, such as Black female underrepresentation and same-race pair preference, survive. My agreement with the reader is 'agree' because we identify the same weakest point; I make the mechanism more specific by pointing to the confidence threshold and the three-race taxonomy.","tokens_in":11708,"tokens_out":8125,"duration_ms":92882,"concrete_test":"Select a stratified random sample of about 2,000 detected faces from the 85,957 images (strata: platform, Face++ race, Face++ gender, and confidence bin, including faces that would be dropped by the 0.7 threshold), obtain consensus manual labels of race, gender, and smile intensity from at least two annotators, and build a confusion matrix for Face++. Use the matrix to re-estimate the Section 4 proportions and the Section 5 pair z-scores, for instance by inverse matrix adjustment or by re-running the pipeline on only high-agreement samples. If the adjusted Black female proportion stays near 3% and the median same-race Z-scores stay positive on both platforms, the concern is resolved; otherwise the headline numbers need to be revised or re-scoped as Face++-relative measurements.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Every headline number in Sections 4–6 is a statement about Face++ outputs, not directly about people in the ads, and the paper provides no validation of Face++ on this corpus. Section 3 sets a confidence threshold of 0.7; if low-confidence faces are not uniformly distributed across skin tones and genders, the sample is selected in a way that can mimic or mask underrepresentation before any label error enters. Section 7 cites Gender Shades [13] to concede 'inherent biases,' but that admission is never quantified here. In particular, Face++'s three-race output (White/Asian/Black) forces Hispanic, Middle Eastern, and multiracial models into those categories or drops them, directly affecting the census comparison in Section 4. A systematic tendency to misclassify Black women would change the 2.9–3.6% Black female figure and could change the same-race z-scores in Figure 4(b). Because the data and code are not released, no reader can assess the magnitude of these risks.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an automated pipeline for measuring gender and racial diversity in advertising images on social media, and demonstrates it on 85,957 images posted by 73 U.S.-originated B2C brands on Instagram and Facebook. Using Face++ for face detection and demographic inference, the authors report the frequencies of six inferred demographic groups (Asian/Black/White by female/male), brand-level gender and racial diversity indices, an analysis of cross-sex interaction pairs with a permutation null model, and an analysis of smile levels. The headline findings are that White models dominate (74.2% of labeled faces on Instagram and 79.3% on Facebook), Black female models appear at only 2.9% and 3.6% of labeled faces on the two platforms, Black models are underrepresented relative to the U.S. Census while Asian models are overrepresented, same-race male-female pairs are preferred in cross-sex interactions once base rates are accounted for, and female models smile more than male models. The paper frames the contribution as a first step toward a fully automated watchdog for diversity in online advertising.","tokens_in":11886,"tokens_out":6235,"duration_ms":66722,"significance":"If the measurement pipeline were validated for this corpus, the paper would provide a valuable large-scale extension of content-analytic studies of advertising from TV and print to social media, with three concrete, computable diversity metrics and a permutation-based null model for pair preferences. The authors are careful to use typewriter font to denote Face++-inferred groups, and they acknowledge in Section 7 that the face-detection tool may have inherent biases. However, the validity of every headline number in Sections 4-6 rests on the accuracy of Face++ for this specific set of advertising images, and that accuracy is not established. The census comparison in Section 4 also conflates Face++'s three race categories with U.S. Census categories, which is a category-mismatch problem rather than a mere labeling-error problem. The contribution is therefore best seen as a promising methodological demonstration whose quantitative conclusions require additional validation and sensitivity analysis.","major_comments":[{"comment":"The central descriptive claims—White dominance, Black female underrepresentation at 2.9-3.6%, the race-census comparison, the same-race z-scores in Fig. 4(b), and the smile-level results—are all statements about Face++ outputs, yet Face++ is never validated on this corpus. The only accuracy evidence cited is the authors' own prior study (ref [27]) on different data, and the manuscript itself cites Gender Shades (ref [13]) which documents systematic error disparities for darker-skinned women. In addition, the confidence threshold of 0.7 in Section 3 may exclude faces non-uniformly by skin tone, gender, or age, introducing selection bias before any label error is considered. The Section 7 acknowledgement that 'the use of Face detection ... may have inherent biases' is not quantified or addressed empirically. I request a stratified validation sample (by platform, inferred group, and confidence score) with human labels, reported as confusion matrices, plus a sensitivity analysis showing how the headline proportions and z-scores change under plausible misclassification rates. Without this, the load-bearing quantitative claims are not supported.","section":"Section 3; Sections 4-6"},{"comment":"The comparison of Face++ race proportions to the 2010 U.S. Census is not valid as stated. Face++ outputs exactly three race categories (White, Asian, Black), so Hispanic, Middle Eastern, and multiracial models must be forced into one of these or dropped, while the Census treats Hispanic as an ethnicity that can co-occur with any race. The normalized census percentages (5.6% Asian, 14.4% Black, 80% White) are therefore not a proper baseline for 'overrepresented' or 'underrepresented' claims. For example, a Hispanic model may be labeled White by Face++, inflating White counts, and the Census 'White alone' population includes Hispanic Whites. The claims should be reframed as relative to a baseline defined by the same three categories, or the authors should validate a subset with human labels into those exact three categories and compare to that. This directly affects RQ1 and the conclusion that Black models are underrepresented relative to the U.S. population.","section":"Section 4, 'Compared with the actual population' paragraph"},{"comment":"The permutation null model is conceptually reasonable, but the manuscript does not specify whether the shuffle is performed separately within each platform and brand, nor how pairs are counted in the null when an image contains more than two models. More importantly, the z-score formula assumes approximate normality of the null distribution of pair counts; for rare pairs such as A/F-B/M, the distribution will be highly skewed and possibly zero-inflated, making the z-score and its boxplots in Fig. 4(b) hard to interpret. The claim that same-race pairs are preferred relies on these z-scores, so I ask the authors to provide details of the shuffling procedure, report null distributions or exact p-values/confidence intervals for rare pairs, and show that the same-race preference is robust to alternative null models (e.g., preserving race as well as gender composition per image).","section":"Section 5, 'To compute the corrected preference' paragraph"}],"minor_comments":[{"comment":"There are several typos and formatting issues: 'ANOV A' should be 'ANOVA'; in Section 6 '55.6 ad.nd 42.8' should be '55.6 and 42.8'; the subscript formatting for N^b_Female and N^b_Male in Section 4 is broken.","section":"Throughout"},{"comment":"Figure 2 is extremely dense, with over 70 brand labels and two connected markers per brand; the labels overlap substantially. A zoomable version or a faceted plot would improve readability.","section":"Figure 2"},{"comment":"The notation 'A/', 'B/', 'W/' is only explained in the figure caption; consider defining it in the main text at first use.","section":"Section 5, Figure 3 caption"},{"comment":"The t-test and Kruskal-Wallis tests treat faces as independent units, but faces are clustered within images and brands. This likely inflates significance; a mixed-effects model or cluster-robust standard errors would be more appropriate.","section":"Section 6, statistical tests"},{"comment":"The paper does not release the image-level data or code, which limits reproducibility of the pipeline. At minimum, a detailed protocol and the list of brand account handles should be provided.","section":"Section 7, limitations"}],"recommendation":"major_revision","confidential_remarks":"The paper's novelty relative to the authors' earlier SocInfo paper (ref [7]) should be clarified; the current manuscript appears to extend that work from 69 to 73 brands and adds the cross-sex interaction and smile analyses, but the framing as a new 'first step' may overstate novelty. Additionally, the central accuracy justification for Face++ relies on a self-authored validation study (ref [27]); given the paper's reliance on Face++ outputs, independent validation on the actual corpus is essential. The census comparison issue is serious enough that it should be addressed before publication, not merely in a limitations paragraph."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a competent, clearly written descriptive study that extends the authors' earlier 69-brand analysis to 73 US-origin brands and adds two new dimensions: cross-sex interaction pairings with a permutation null model, and smile-level analysis. The same-race preference result in Figure 4(b) is the real new finding and it is handled carefully: they shuffle model placement while preserving per-image gender counts, and the z-scores are an appropriate correction for base rates. The smile findings replicate a well-known gender difference (women smile more) and add race breakdowns.\n\nThe soft spots are real but not fatal. First, the paper calls the sample 'international' but only includes US-origin brands, as the data section actually states; the abstract oversells. Second, every proportion in Sections 4-6 is a statement about Face++ outputs, not directly about people in the ads. The authors cite Gender Shades but do not validate Face++ on this corpus, and the 0.7 confidence threshold could differentially drop low-confidence faces. The Black female underrepresentation numbers (2.9-3.6%) and the same-race z-scores could shift with classifier error. Third, the leap from same-race pairings to 'implicit opposition to interracial relationships' goes beyond what the data can support; the cited survey evidence is about attitudes, not ad production decisions. Fourth, data and code are not released, so the analysis cannot be independently checked.\n\nWhat the paper does well: the permutation null model is a genuine improvement over raw proportions, the brand-level variance analysis is informative, and the writing is transparent about limitations. The citation pattern is fine; the self-citations are to the actual precursor and the validation study, which is legitimate, though the validation study (ref 27) is the authors' own and not independent.\n\nBottom line: as a description of what Face++ sees in these ads, the results are probably reliable, and the same-race pairing result is worth reporting. As a statement about actual diversity in social media advertising, it needs a stratified validation of the classifier or explicit error bars. This deserves peer review—it is a solid empirical contribution with a clear method and a new result—but it should go back for a revision that either validates Face++ on a labeled sample of these images or substantially softens the headline claims. I'd bring it to a reading group focused on computational social science or media studies, and I'd cite it (with caveats) if I were working on automated advertisement auditing.","headline":"A solid descriptive extension of the authors' own 2018 study, with a genuinely new permutation-based result on same-race cross-sex pairings, but every headline number depends on unvalidated Face++ labels.","tokens_in":12393,"tokens_out":1529,"would_cite":true,"duration_ms":15876,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Automated scan of 85,957 brand ads finds White models dominate and same-race pairings are favored.","keywords":["gender diversity","racial diversity","advertising images","social media advertising","face detection","Face++","cross-sex interactions","smiling"],"falsifier":"Take a stratified random sample of the 85,957 images, have independent human annotators label gender and race for each detected face, and compare those labels with Face++'s output; if human labels put Black female models near their 14.4% population share while Face++ put them near 2.9%, or if human-labeled smile scores show no female–male gap, the paper's headline numbers would not survive.","tokens_in":11479,"feed_emoji":"📊","tokens_out":9588,"duration_ms":91087,"temperature":0.7,"pith_summary":"The paper tries to show that a fully automated image-analysis pipeline can measure gender and racial diversity in online brand advertising, and that applying it to 85,957 Instagram and Facebook images from 73 top U.S.-origin consumer brands reveals systematic patterns: White models dominate (74.2% of labeled faces on Instagram, 79.3% on Facebook), Black models are underrepresented relative to the U.S. population, Asian models are overrepresented, and after correcting for base rates with a permutation null model, same-race male–female pairs are favored in cross-sex interactions. It also reports that female models smile significantly more than male models. The authors care because repeated exposure to mediated images can shape perceptions of demographic groups, and previous monitoring has relied on slow manual content analysis. If the approach works, diversity in online ads could be tracked cheaply and near real time.","feed_headline":"White models dominate 85,957 social media brand ads","feed_subtitle":"A fully automated audit finds Black models at half their U.S. population share and same-race couples favored in mixed-gender ads.","key_machinery":"The mechanism is a fully automated pipeline built on the commercial face-analysis service Face++, which detects faces and infers gender (Female or Male), race (White, Asian, or Black), and a smile score from 0 to 100; images with confidence below 0.7 are discarded. On top of these labels, the paper defines three computable metrics: appearance frequency for each demographic group, pair frequencies for cross-sex interactions (images containing at least one male and one female model), and smile-level distributions. The load-bearing statistical object is the permutation null model used for cross-sex pairs: it shuffles models across images while keeping the number of male and female models in every image fixed, then computes 1,000 randomized replicates to produce a z-score for each of the nine race-by-gender pairs. That z-score is what separates a pair appearing often because those models appear often from a pair appearing more often than the base rates would predict.","core_discovery":"The central discovery is that the demographic imbalances long documented in TV and print advertising reappear in social media advertising, and that at least one imbalance—same-race pairing in cross-sex interactions—is not just a byproduct of White models being more numerous. On the raw counts, White female–White male pairs dominate cross-sex ads (60.3% of pairs on Instagram, 68.9% on Facebook). But the paper's permutation null model, which shuffles models across images while preserving the number of men and women in each image, shows that same-race pairs of every race have positive median z-scores on both platforms, while interracial pairs involving Whites have negative z-scores. In other words, after accounting for how often each group appears, advertisers still pair models of the same race more often than chance. The paper also claims that Black female models are rare (2.9% of Instagram models, 3.6% on Facebook) and that female models smile more than male models with a large effect size.","pith_inferences":["A natural extension is to rerun the same pair analysis on images with exactly two models, where the romantic or family cue is strongest; the paper pools all pairs, so multi-model images may dilute or exaggerate the same-race preference.","The roughly threefold overrepresentation of Asian models could be tested against regional marketing motives by comparing accounts targeting U.S. audiences with those targeting Asia; the paper uses only primary accounts of U.S.-origin brands.","The pipeline's confidence threshold could itself be a source of bias: if Face++ drops low-confidence faces unevenly across demographic groups, the reported proportions would be distorted, and that is checkable by comparing the demographic mix of accepted and rejected faces.","An automated watchdog built on this approach could add other visual attributes—age, body type, clothing, or the role a model plays—to catch stereotyped portrayals that simple headcount ratios miss; the authors note the need for such nuance in their limitations."],"forward_implications":["Brand diversity can be tracked continuously and cheaply: the same pipeline can be rerun as new images are posted, turning diversity monitoring into a near-real-time audit.","The specific ratios (White 74.2% on Instagram and 79.3% on Facebook; Black 8.9% and 7.4% versus a normalized U.S. population share of 14.4%) give advertisers and watchdogs a concrete, reproducible baseline to measure against.","Because same-race pairing survives the permutation correction, social media ads join TV and print as a site where interracial romantic and family cues are underused.","The large smile gap between female and male models (Cohen's d is about 1.05) indicates that gendered display norms persist in online brand imagery, not just in traditional media."],"supporting_citations":[{"why":"supplies the Face++ face detector and the gender, race, and smile labels used on all 85,957 images.","marker":"[2]"},{"why":"provides the 2010 U.S. Census population shares that serve as the benchmark for calling Black models underrepresented and Asian models overrepresented.","marker":"[5]"},{"why":"is the accuracy study the authors cite for Face++ achieving over 90% accuracy in gender and race inference.","marker":"[27]"},{"why":"documents White preference in cross-sex interactions in 1990s TV advertising, the pattern the paper extends to social media.","marker":"[16]"},{"why":"is the study showing commercial gender classifiers can be substantially less accurate for darker-skinned women, which the authors acknowledge as a limitation of the labeling step.","marker":"[13]"},{"why":"is the review of decades of broadcast and print advertising portrayals that the paper compares with its social-media findings, particularly for Black models.","marker":"[46]"},{"why":"is the authors' earlier social-media diversity case study that this work expands to more brands and to the cross-sex interaction and smile analyses.","marker":"[7]"}],"fun_headline_variants":["Same-race couples over-represented in brand ads even after size control","Black models appear at half their US population share in 86K brand ads","Same-race pair bias in cross-sex ads persists after permutation test","Female models smile more, same-race pairing persists in social ads","Automated audit of 86K social ads finds racial pairing bias beyond headcount"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"If Face++'s gender, race, and smile labels are systematically wrong for these advertising images—especially for darker-skinned women—every reported proportion and pairing preference could be an artifact.","fun_headline_variants_meta":{"raw":{"variants":["Same-race couples over-represented in brand ads even after size control","Black models appear at half their US population share in 86K brand ads","Same-race pair bias in cross-sex ads persists after permutation test","Female models smile more, same-race pairing persists in social ads","Automated audit of 86K social ads finds racial pairing bias beyond headcount"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000607,"raw_usage":{"total_tokens":2755,"prompt_tokens":797,"completion_tokens":1958,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":413,"completion_tokens_details":{"reasoning_tokens":1871}},"tokens_in":413,"tokens_out":1958,"duration_ms":15017,"temperature":1.0,"reasoning_tokens":1871,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:15:51.189075+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a stratified random sample of the 85,957 images, have independent human annotators label gender and race for each detected face, and compare those labels with Face++'s output; if human labels put Black female models near their 14.4% population share while Face++ put them near 2.9%, or if human-labeled smile scores show no female–male gap, the paper's headline numbers would not survive.","supporting_citations":[{"cited_title":"https://www.faceplusplus.com, accessed: 2019-04-16","cited_arxiv_id":null,"evidence_quote":"supplies the Face++ face detector and the gender, race, and smile labels used on all 85,957 images."},{"cited_title":"https://www.census.gov/prod/cen2010/ briefs/c2010br-14.pdf, accessed: 2019-04-16","cited_arxiv_id":null,"evidence_quote":"provides the 2010 U.S. Census population shares that serve as the benchmark for calling Black models underrepresented and Asian models overrepresented."},{"cited_title":"In: ICWSM (2018)","cited_arxiv_id":null,"evidence_quote":"is the accuracy study the authors cite for Face++ achieving over 90% accuracy in gender and race inference."},{"cited_title":"Sex roles 42(5), 363–389 (2000)","cited_arxiv_id":null,"evidence_quote":"documents White preference in cross-sex interactions in 1990s TV advertising, the pattern the paper extends to social media."},{"cited_title":"In: Conference on fairness, accountability and transparency","cited_arxiv_id":null,"evidence_quote":"is the study showing commercial gender classifiers can be substantially less accurate for darker-skinned women, which the authors acknowledge as a limitation of the labeling step."},{"cited_title":"Journal of business ethics 123(3), 421–436 (2014)","cited_arxiv_id":null,"evidence_quote":"is the review of decades of broadcast and print advertising portrayals that the paper compares with its social-media findings, particularly for Black models."},{"cited_title":"In: International Conference on Social Informatics","cited_arxiv_id":null,"evidence_quote":"is the authors' earlier social-media diversity case study that this work expands to more brands and to the cross-sex interaction and smile analyses."}],"review_version":1}