Pith. sign in

REVIEW 3 major objections 7 minor 67 references

Learning AI Auditing: A Case Study of Teenagers Auditing a Generative AI Model

T0 review · 3 major / 7 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read With light scaffolding—picking a hypothesis, designing prompts, running 1,200 tests, and reporting—teenagers can conduct full audits of a real generative AI system, and the paper shows their conclusions align with expert re-analysis.

desk verdict Genuinely new step in teen-led auditing, with a triangulation claim that is suggestive rather than conclusive. read the letter →

arxiv 2508.04902 v2 pith:A6WI7XDD submitted 2025-08-06 cs.HC cs.CY

classification cs.HCcs.CY
keywords algorithmauditingAIliteracyteenagersparticipatorydesigngenerativealgorithmicbiasTikTokEffectHousecasestudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that high-school-age teenagers, given a light scaffold, can conduct a complete algorithm audit of a generative AI system they actually use—deciding what to test, generating inputs, running tests, analyzing results, and reporting findings—and that their conclusions are credible. In a two-week workshop, 14 teens aged 14–15 audited the generative image model behind TikTok's Effect House by testing 25 occupations across four prompts and 12 input images, producing 1,200 outputs. When the research team re-analyzed the same data with expert audit methods, the directions of bias matched the teens' findings (e.g., masculine-presenting outputs for 'carpenter' and 'rapper', feminine for 'nail technician' and 'fast food worker'), even though exact percentages differed. Teens also independently introduced age-related bias, a dimension rarely covered in expert audits. If true, this means audit-based activities could serve as AI literacy education and give young people a concrete role in holding the systems they use to account.

What carries the argument

The five-step algorithm audit process—developing a hypothesis, generating inputs, running tests, analyzing results, and reporting findings—is the scaffolding mechanism that carries the argument. The paper takes this expert methodology (as established in prior auditing research) and transfers it to a participatory workshop, with teens as the protagonists. The audit's empirical engine is the structured test matrix: 25 occupations × 4 prompts × 12 input images = 1,200 query-output pairs, which provides enough data for both teen analysis and the researchers' subsequent expert coding, the latter using categories such as change in gender representation, change in racial representation, wrinkles, a

What would settle it

Rerun the same audit protocol with an independent team of expert coders using the same categories, and compare their labels with the original researchers' labels on the same 240-image overlap; if inter-coder agreement drops or the teen-expert directional agreement disappears, the conclusion that the workshop produced credible audit findings would be undercut. Alternatively, rerun the audit at a later date on a fresh sample from the same Effect House model and check whether the bias patterns replicate.

Watch

Extended reading notes

Core claim

The paper claims that, with appropriate scaffolding, teenagers can participate in full-fledged audits of real-world algorithmic systems they encounter daily, and that a five-step workshop design (hypothesis, input generation, testing, analysis, reporting) can support them in reaching evidence-based, credible conclusions. In the reported case, the teens themselves chose to investigate whether Effect House's generative model reinforces gender and race stereotypes about occupations, generated a list of occupations and prompts, ran 1,200 tests collaboratively, analyzed subsets of the data with varied strategies (perceptual labeling, feature-based counting, comparative description), and created a

Load-bearing premise

The claim that the teens reached 'credible' conclusions is judged against the researchers' own coding of the same images, so it stands only if those expert judgments reliably capture the model's real biases and if 1,200 tests on a proprietary, potentially changing model are representative enough for system-wide conclusions.

Editorial extensions

If this is right

  • The five-step audit scaffold can be reused by educators to build AI literacy activities that move beyond discussion into empirical investigation of real systems.
  • Teen-led audits can surface bias dimensions—notably age—that professional audits often overlook, suggesting that diversity of auditors is not just a nicety but a source of new findings.
  • If audit-based learning is adopted in classrooms, tools that streamline annotation and provide pre-annotated datasets will be needed, since teens could not annotate all 1,200 tests in the workshop.
  • The finding that teens' conclusions track expert triangulation supports the credibility of participatory and end-user auditing more broadly, not just for youth.
  • Audit results targeted at developers, filter designers, and peer users suggest a communication pathway that could make algorithmic harms more visible to the people who can address them.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The teens' choice of novel occupations (rapper, nail technician) and their attention to age suggest that auditors who are themselves heavy users of a platform may systematically select probes different from those chosen by expert auditors; this could be formalized by comparing occupational category distributions between teen-generated and BLS-derived input sets.
  • The paper's credibility argument rests on expert coding as ground truth, but the paper honestly reports initial low inter-rater agreement on skin complexion and gender exaggeration, leading to coarser categories; a natural extension is to measure whether teen-expert agreement persists when the expert ground truth is itself unstable or contested.
  • The 1,200-test dataset was collected from a proprietary, potentially changing model at one point in time; a testable extension is to rerun the same audit protocol after a model update to see whether teen-audited bias patterns are stable or drift, which would inform how often participatory audits need to be repeated.
  • The 'sixth step' the authors muse about—having teens reflect on the audit process—could plausibly increase the durability of the AI literacy gains, but the paper does not test this; a pre/post reflection-condition comparison would be needed.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper reports a descriptive case study of a two-week participatory design workshop in which 14 teenagers (ages 14-15) audited the generative image model behind TikTok's Effect House. After exploratory play, the group converged on the hypothesis that the model reinforces gender and race stereotypes about occupations, collaboratively generated 25 occupation prompts and 12 input images, ran 1,200 tests, and analyzed a subset of 13 occupations in small groups using heterogeneous coding schemes (perceived gender/race/age, concrete attributes like wrinkles and facial hair, or qualitative descriptions). The authors then triangulated the teens' findings by having three expert researchers code the full 1,200-image dataset on gender, age, and race, reporting Fleiss's kappa = 0.65 and qualitative/numeric agreement on several occupations. The paper claims that the workshop enabled teenagers to conduct full-fledged audits and to reach evidence-based, credible conclusions, while also surfacing novel dimensions such as age bias. The main research questions are procedural engagement and credibility of conclusions.

Significance. If the triangulation claim is established, this is a valuable contribution to participatory auditing, AI literacy, and child-computer interaction. The study is among the first to engage teenagers in an end-to-end audit of a real, widely used generative AI system, and the workshop design is described in enough detail to be adapted. The authors are transparent about the instability of their initial expert codes, provide a complete coding scheme in Appendix A, and explicitly acknowledge the single-case nature of the study in Section 5.4. The youth's independent identification of age-related bias is a genuinely novel and interesting finding for the algorithm-auditing literature. However, the central empirical evidence for 'evidence-based, credible conclusions' rests on a small set of selected examples rather than a systematic comparison, and the expert ground truth itself has documented reliability concerns. These issues are fixable but require additional analysis or more carefully qualified claims.

major comments (3)
  1. [§4.2 and Table 3] The paper's strongest claim—that the workshop supported teens in reaching 'evidence-based, credible conclusions'—rests on the triangulation in §4.2, but the evidence presented is a set of favorable examples rather than a complete mapping. Teens analyzed 13 of the 25 occupations (§4.1.4) with heterogeneous, group-specific coding schemes, while the expert analysis covered all 25 occupations. The text does not enumerate agreements, disagreements, and non-comparable claims occupation by occupation. In at least one case the comparison uses different constructs: Ibrahim's finding that 100% of rapper outputs had 'darkened skin' is compared with an expert finding of 60% 'change in racial representation,' which is not the same measure. To make the triangulation claim load-bearing, the authors should provide a systematic table mapping each teen group finding to the corresponding expert-coded resul
  2. [§3.5.2] The expert ground truth used for triangulation is not as stable as the summary suggests. Initial agreement on skin complexion was 50.42% and on gender exaggeration 50.83%, which prompted the authors to replace those codes with coarser categories ('change in gender/racial representation'). This replacement is reasonable, but it changes the construct being validated and may improve reliability simply because coarser categories are easier to agree on. The reported overall Fleiss's kappa of 0.65 (95% CI 0.58-0.72) is an average across all categories and can obscure low reliability on specific codes. The paper should report per-code reliabilities for the final coding scheme and, if possible, show that the final codes are not so coarse that they trivialize the comparison. Because the entire RQ2 answer depends on expert coding, this is a load-bearing point.
  3. [§3.5.2 and §4.2 (age analysis)] The age dimension was added to the expert coding only after the teens reported age-related findings ('we also added age representation since several groups of participants also analyzed this kind of bias'). Using that same post hoc expert coding to validate the teens' age findings is not independent confirmation. The age codes themselves (wrinkles, gray hair) are concrete and defensible, but the paper should explicitly acknowledge that for age the expert analysis was constructed after seeing the teen findings, and therefore the agreement on age is weaker evidence of credibility than the agreement on gender/race. If the authors have any pre-registered or earlier analysis plan that included age, that should be stated; otherwise the claim that youth 'independently raising' age bias is a contribution should be separated from the claim that expert analysis confirmed it.
minor comments (7)
  1. [Table 1] The row for Day 1 contains a typo: 'led a an activity' should be 'led an activity.'
  2. [§3.5.2] 'perspectices' should be 'perspectives.'
  3. [Figure 3 and §4.1.1] The name is spelled inconsistently: 'Kaden' in the Figure 3 caption but 'Kayden' in the text and Table 3. Please standardize.
  4. [§4.1.3 and Figure 4] 'Jason Mamoa' is misspelled; the correct spelling used elsewhere is 'Jason Momoa.'
  5. [Table 2] 'New anchor' should be 'News anchor.'
  6. [§4.1.2] The parenthetical '(see Figure 2' appears to be a broken cross-reference; it should likely refer to Table 2 (the occupations list), and the closing parenthesis is missing.
  7. [§3.3] The text says 12 input images were used, but the description lists 4 Effect House defaults plus 10 celebrity images = 14, with images 13 and 14 not used. Consider clarifying this arithmetic in one sentence.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the teen-expert triangulation is an independent re-analysis of the same dataset, not an input-derived prediction.

full rationale

The paper's central claim—that the workshop supported teens in reaching evidence-based, credible conclusions—is checked against the authors' own re-coding of the same 1200-image dataset. This re-coding is not derived from the teens' findings: gender and race were part of the original hypothesis, and the coding scheme uses observable proxies (e.g., facial hair, wrinkles, gray hair, perceived gender/racial change) applied independently by three researchers. The age dimension was added post hoc because teens analyzed it, which is a methodological selection choice rather than a logical reduction: the age coding was still performed independently and could in principle have contradicted the teens' age-related claims. The paper's self-citations (e.g., Metaxa et al., Morales-Navarro et al.) are used as background methodology and prior work, not as a uniqueness argument or to force the findings. Some comparisons between teen and expert numbers use different constructs (e.g., 'darkened skin' vs. 'change in racial representation'), but that is a measurement-validity concern, not circularity. No equation or fitted parameter is renamed as a prediction, and no derivation reduces by construction to its own inputs. Therefore, under the strict standard requiring a specific reduction, no significant circularity is present.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central claims rest on methodological assumptions about the validity of subjective coding as ground truth and the representativeness of the test set; no new entities are introduced, and the only hand-picked numeric parameter is the Effect House generation setting controlled by the researchers.

free parameters (1)
  • Effect House generation settings (prompt strength, style, transition style) = Prompt Strength 0.50, Style Stylized II, Transition Style Style 1
    Chosen by researchers to standardize tests (Section 3.3); Horacio's re-testing at strength 1.0 showed the setting materially affects gender change observations (Section 4.1.4), so the audit conclusions depend on this hand-picked value.
assumptions (4)
  • domain assumption Expert coding of gender, race, and age changes is a valid measure of bias in the generated images.
    Used throughout Section 3.5.2 and Section 4.2 as the ground truth for triangulation; initial inter-rater agreement was low on several categories, making this assumption fragile.
  • domain assumption The 1,200 tests constitute a representative sample of Effect House's model behavior.
    Section 3.3 relies on this to draw system-wide conclusions; the model is proprietary and may vary over time, and inputs were non-random.
  • domain assumption Researcher-facilitators did not lead participants to predetermined conclusions during scaffolding.
    Section 3.3 shows a researcher proposed the final hypothesis after teens brainstormed, which could shape the audit's direction and the subsequent triangulation.
  • domain assumption Self-reported engagement and screen-recording observations accurately reflect what teens learned.
    Findings in Section 4.1 are based on qualitative observations and videologs, which cannot be independently verified from the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Learning AI Auditing: A Case Study of Teenagers Auditing a Generative AI Model." pith.science (2026). https://pith.science/paper/A6WI7XDD

@misc{pith2026250804902,
  author       = {Pith},
  title        = {Pith review of: Learning AI Auditing: A Case Study of Teenagers Auditing a Generative AI Model},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/A6WI7XDD}},
  note         = {Machine review of arXiv:2508.04902}
}
read the original abstract

This study investigates how high school-aged youth engage in algorithm auditing to identify and understand biases in artificial intelligence and machine learning (AI/ML) tools they encounter daily. With AI/ML technologies being increasingly integrated into young people's lives, there is an urgent need to equip teenagers with AI literacies that build both technical knowledge and awareness of social impacts. Algorithm audits (also called AI audits) have traditionally been employed by experts to assess potential harmful biases, but recent research suggests that non-expert users can also participate productively in auditing. We conducted a two-week participatory design workshop with 14 teenagers (ages 14-15), where they audited the generative AI model behind TikTok's Effect House, a tool for creating interactive TikTok filters. We present a case study describing how teenagers approached the audit, from deciding what to audit to analyzing data using diverse strategies and communicating their results. Our findings show that participants were engaged and creative throughout the activities, independently raising and exploring new considerations, such as age-related biases, that are uncommon in professional audits. We drew on our expertise in algorithm auditing to triangulate their findings as a way to examine if the workshop supported participants to reach coherent conclusions in their audit. Although the resulting number of changes in race, gender, and age representation uncovered by the teens were slightly different from ours, we reached similar conclusions. This study highlights the potential for auditing to inspire learning activities to foster AI literacies, empower teenagers to critically examine AI systems, and contribute fresh perspectives to the study of algorithmic harms.

Figures

Figures reproduced from arXiv: 2508.04902 by the authors.

Figure 1
Figure 1. Example of a TikTok filter in action. This filter uses generative AI to output a manga-style illustration [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Effect House’s visual scripting, prompting, and filter preview interfaces. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Inputs and outputs for Kaden’s experiments with the prompts “tennis player” and “basketball player.” [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: The input images used by participants in their audit, including 4 default Effect House inputs and 10 [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]
Figure 5
Figure 5. Figure 5: Screenshot of a section of the spreadsheet where participants kept track of their tests. [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 6
Figure 6. Figure 6: Examples of input and output images for three prompts. [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]
Figure 7
Figure 7. Figure 7: Bar graphs depicting the gender representations (left graph) and the changes in racial representations [PITH_FULL_IMAGE:figures/full_fig_p019_7.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

67 extracted references · 44 canonical work pages

  1. [1]

    Bobby Allyn, Sylvia Goodman, and Dara Kerr. 2024. TikTok executives know about app’s effect on teens, lawsuit documents allege. https://www.npr.org/2024/10/11/g-s1-27676/tiktok-redacted-documents-in-teen-safety-lawsuit- revealed

  2. [2]

    Monica Anderson, Michelle Faverio, and Jeffrey Gottfried. 2023. Teens, Social Media and Technology 2023. Pew Research Center. https://www.pewresearch.org/internet/2023/12/11/teens-social-media-and-technology-2023/

  3. [3]

    Jack Bandy. 2021. Problematic Machine Behavior: A Systematic Literature Review of Algorithm Audits. Proc. ACM Hum.-Comput. Interact. 5, CSCW1, Article 74 (April 2021), 34 pages. doi:10.1145/3449148

  4. [4]

    Liam J Bannon and Pelle Ehn. 2012. Design: design matters in Participatory Design. InRoutledge international handbook of participatory design. Routledge, New York, 37–63

  5. [5]

    Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, and Aylin Caliskan. 2023. Easily Accessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency (Chicago, IL,...

  6. [6]

    Abeba Birhane. 2021. Algorithmic injustice: a relational ethics approach. Patterns 2, 2 (2021), 1–9

  7. [7]

    Jolie Bonner, Florian Mathis, Joseph O’Hagan, and Mark Mcgill. 2023. When Filters Escape the Smartphone: Exploring Acceptance and Concerns Regarding Augmented Expression of Social Identity for Everyday AR. In Proceedings of the 29th ACM Symposium on Virtual Reality Software and Technology (Christchurch, New Zealand) (VRST ’23). Association for Computing M...

  8. [8]

    Claus Bossen, Christian Dindler, and Ole Sejer Iversen. 2016. Evaluation in participatory design: a literature survey. In Proceedings of the 14th Participatory Design Conference: Full Papers - Volume 1 (Aarhus, Denmark) (PDC ’16). Association for Computing Machinery, New York, NY, USA, 151–160. doi:10.1145/2940299.2940303

Show all 67 references
  1. [9]

    Joy Buolamwini and Timnit Gebru. 2018. Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. In Proceedings of the 1st Conference on Fairness, Accountability and Transparency (Proceedings of Machine Learning Research, Vol. 81), Sorelle A. Frie...

  2. [10]

    Kaitlyn Burnell, Allycen R Kurup, and Marion K Underwood. 2022. Snapchat Lenses and Body Image Concerns. New Media & Society 24, 9 (2022), 2088–2106. doi:10.1177/1461444821993038

  3. [11]

    That’s what techquity is

    Merijke Coenraad. 2022. “That’s what techquity is”: youth perceptions of technological and algorithmic bias.Information and Learning Sciences 123, 7/8 (2022), 500–525

  4. [12]

    Norman K Denzin. 2017. The research act: A theoretical introduction to sociological methods . Routledge, New York

  5. [13]

    Alicia DeVrio, Aditi Dhabalia, Hong Shen, Kenneth Holstein, and Motahhare Eslami. 2022. Toward User-Driven Algorithm Auditing: Investigating users’ strategies for uncovering harmful algorithmic behavior. In Proceedings of the 2022 CHI Conference on Human Factors in Computing S...

  6. [14]

    Christian Dindler, Ole Sejer Iversen, Mikkel Hjorth, Rachel Charlotte Smith, and Hannah Djurssø Nielsen. 2023. DORIT: An analytical model for computational empowerment in K-9 education. International Journal of Child-Computer Interaction 37 (2023), 100599

  7. [15]

    Miriam Doh, Corinna Canali, and Anastasia Karagianni. 2024. Pixels of Perfection and Self-Perception: Deconstructing AR Beauty Filters and Their Challenge to Unbiased Body Image. InProceedings of the 2024 ACM International Conference on Interactive Media Experiences (Stockholm...

  8. [16]

    Ana Maria Bustamante Duarte, Nina Brendel, Auriol Degbelo, and Christian Kray. 2018. Participatory Design and Participatory Research: An HCI Case Study with Young Forced Migrants. ACM Trans. Comput.-Hum. Interact. 25, 1, Article 3 (Feb. 2018), 39 pages. doi:10.1145/3145472

  9. [17]

    Eureka Foong, Darren Gergle, and Elizabeth M. Gerber. 2017. Novice and Expert Sensemaking of Crowdsourced Design Feedback. Proc. ACM Hum.-Comput. Interact. 1, CSCW, Article 45 (Dec. 2017), 18 pages. doi:10.1145/3134680

  10. [19]

    Ole Sejer Iversen, Rachel Charlotte Smith, and Christian Dindler. 2017. Child as Protagonist: Expanding the Role of Children in Participatory Design. In Proceedings of the 2017 Conference on Interaction Design and Children (Stanford, California, USA) (IDC ’17). Association for...

  11. [20]

    Niharika Jain, Alberto Olmo, Sailik Sengupta, Lydia Manikonda, and Subbarao Kambhampati. 2022. Imperfect ImaGANation: Implications of GANs exacerbating biases on facial data augmentation and snapchat face lenses. Artificial Intelligence 304 (2022), 103652

  12. [21]

    Nadia Karizat, Dan Delmonaco, Motahhare Eslami, and Nazanin Andalibi. 2021. Algorithmic Folk Theories and Identity: How TikTok Users Co-Produce Knowledge of Identity and Engage in Algorithmic Resistance. Proceedings of the ACM on Human-Computer Interaction 5 (2021), 305:1–305:...

  13. [22]

    Matthew Kay, Cynthia Matuszek, and Sean A. Munson. 2015. Unequal Representation and Gender Stereotypes in Image Search Results for Occupations. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (Seoul, Republic of Korea) (CHI ’15). Associat...

  14. [23]

    Ko, Alannah Oleson, Neil Ryan, Yim Register, Benjamin Xie, Mina Tari, Matthew Davidson, Stefania Druga, and Dastyni Loksa

    Amy J. Ko, Alannah Oleson, Neil Ryan, Yim Register, Benjamin Xie, Mina Tari, Matthew Davidson, Stefania Druga, and Dastyni Loksa. 2020. It is time for more critical CS education. Commun. ACM 63, 11 (Oct. 2020), 31–33. doi:10. 1145/3424000

  15. [24]

    Lam, Mitchell L

    Michelle S. Lam, Mitchell L. Gordon, Danaé Metaxa, Jeffrey T. Hancock, James A. Landay, and Michael S. Bernstein

  16. [25]

    Lam, Ayush Pandit, Colin H

    Michelle S. Lam, Ayush Pandit, Colin H. Kalicki, Rachit Gupta, Poonam Sahoo, and Danaé Metaxa. 2023. Sociotechnical Audits: Broadening the Algorithm Auditing Lens to Investigate Targeted Advertising. Proc. ACM Hum.-Comput. Interact. 7, CSCW2, Article 360 (Oct. 2023), 37 pages....

  17. [26]

    Melissa R Laughter, Jaclyn B Anderson, Mayra BC Maymone, and George Kroumpouzos. 2023. Psychology of aesthetics: Beauty, social media, and body dysmorphic disorder. Clinics in dermatology 41, 1 (2023), 28–32

  18. [27]

    Lee, Victoria Delaney, and Parth Sarin

    Victor R. Lee, Victoria Delaney, and Parth Sarin. 2022. Eliciting High School Students’ Conceptions and Intuitions about Algorithmic Bias. In Proceedings of the 2022 ACM Conference on International Computing Education Research - Volume 2 (Lugano and Virtual Event, Switzerland)...

  19. [28]

    Ang Li, Zheng Yao, Diyi Yang, Chinmay Kulkarni, Rosta Farzan, and Robert E. Kraut. 2020. Successful Online Socialization: Lessons from the Wikipedia Education Program. Proc. ACM Hum.-Comput. Interact. 4, CSCW1, Article 50 (May 2020), 24 pages. doi:10.1145/3392857

  20. [29]

    Duri Long and Brian Magerko. 2020. What is AI Literacy? Competencies and Design Considerations. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–16. doi:10.1...

  21. [30]

    Sasha Luccioni, Christopher Akiki, Margaret Mitchell, and Yacine Jernite. 2023. Stable Bias: Evaluating Soci- etal Representations in Diffusion Models. In Advances in Neural Information Processing Systems , A. Oh, T. Nau- mann, A. Globerson, K. Saenko, M. Hardt, and S. Levine ...

  22. [31]

    Crawford, Sanjana Gautam, Sorelle A

    Yaaseen Mahomed, Charlie M. Crawford, Sanjana Gautam, Sorelle A. Friedler, and Danaé Metaxa. 2024. Auditing GPT’s Content Moderation Guardrails: Can ChatGPT Write Your Favorite TV Show?. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (R...

  23. [32]

    Harvey Mannering. 2023. Analysing Gender Bias in Text-to-Image Models Using Object Detection. arXiv:2307.08025 [cs] doi:10.48550/arXiv.2307.08025

  24. [33]

    Nora McDonald, Sarita Schoenebeck, and Andrea Forte. 2019. Reliability and Inter-rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice. Proc. ACM Hum.-Comput. Interact. 3, CSCW, Article 72 (Nov. 2019), 23 pages. doi:10.1145/3359174

  25. [34]

    Gan, Su Goh, Jeff Hancock, and James A

    Danaé Metaxa, Michelle A. Gan, Su Goh, Jeff Hancock, and James A. Landay. 2021. An Image of Society: Gender and Racial Representation and Impact in Image Search Results for Occupations. Proc. ACM Hum.-Comput. Interact. 5, CSCW1, Article 26 (April 2021), 23 pages. doi:10.1145/3449100

  26. [35]

    Danaé Metaxa, Joon Sung Park, Ronald E Robertson, Karrie Karahalios, Christo Wilson, Jeff Hancock, Christian Sandvig, et al. 2021. Auditing algorithms: Understanding algorithmic systems from the outside in. Foundations and Trends® in Human–Computer Interaction 14, 4 (2021), 272–344

  27. [36]

    Luis Morales-Navarro, Yasmin Kafai, Vedya Konda, and Danaé Metaxa. 2024. Youth as Peer Auditors: Engaging Teenagers with Algorithm Auditing of Machine Learning Applications. In Proceedings of the 23rd Annual ACM Interaction Design and Children Conference (Delft, Netherlands) (...

  28. [37]

    Luis Morales-Navarro and Yasmin B Kafai. 2023. Conceptualizing approaches to critical computing education: Inquiry, design, and reimagination. In Past, Present and Future of Computing Education Research: A Global Perspective . Springer, Proc. ACM Hum.-Comput. Interact., Vol. 9...

  29. [38]

    Luis Morales-Navarro and Yasmin B Kafai. 2024. Unpacking Approaches to Learning and Teaching Machine Learning in K-12 Education: Transparency, Ethics, and Design Activities. In Proceedings of the 19th WiPSCE Conference on Primary and Secondary Computing Education Research . As...

  30. [39]

    Luis Morales-Navarro, Yasmin B Kafai, Lauren Vogelstein, Evelyn Yu, and Danaé Metaxa. 2025. Learning About Algorithm Auditing in Five Steps: Scaffolding How High School Youth Can Systematically and Critically Evaluate Machine Learning Applications. In Proceedings of the AAAI C...

  31. [40]

    Ranjita Naik and Besmira Nushi. 2023. Social Biases through the Text-to-Image Generation Lens. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society (Montréal, QC, Canada) (AIES ’23). Association for Computing Machinery, New York, NY, USA, 786–808. doi:10.1...

  32. [41]

    Leonardo Nicoletti and Diana Bass. 2023. Humans Are Biased. Generative AI Is Even Worse: Stable Diffusion’s text-to-image model amplifies stereotypes about race and gender — here’s why that matters. https://www.bloomberg. com/graphics/2023-generative-ai-bias/

  33. [42]

    Rizu Paudel and Mahdi Nasrullah Al-Ameen. 2024. Leveraging the Power of Storytelling to Encourage and Empower Children towards Strong Passwords. Proc. ACM Hum.-Comput. Interact. 8, CSCW2, Article 504 (Nov. 2024), 27 pages. doi:10.1145/3687043

  34. [43]

    I Wish I Was Wearing a Filter Right Now

    Claire Kathryn Pescott. 2020. “I Wish I Was Wearing a Filter Right Now”: An Exploration of Identity Formation and Subjectivity of 10- and 11-Year Olds’ Social Media Use. Social Media + Society 6, 4 (2020), 2056305120965155. doi:10.1177/2056305120965155

  35. [44]

    Organizers Of Queerinai, Anaelia Ovalle, Arjun Subramonian, Ashwin Singh, Claas Voelcker, Danica J. Sutherland, Davide Locatelli, Eva Breznik, Filip Klubicka, Hang Yuan, Hetvi J, Huan Zhang, Jaidev Shriram, Kruno Lehman, Luca Soldaini, Maarten Sap, Marc Peter Deisenroth, Maria...

  36. [45]

    Mark S Reed, Bethann Garramon Merkle, Elizabeth J Cook, Caitlin Hafferty, Adam P Hejnowicz, Richard Holliman, Ian D Marder, Ursula Pool, Christopher M Raymond, Kenneth E Wallen, et al . 2024. Reimagining the language of engagement in a post-stakeholder world. Sustainability Sc...

  37. [46]

    Jean Salac, Alannah Oleson, Lena Armstrong, Audrey Le Meur, and Amy J. Ko. 2023. Funds of Knowledge used by Adolescents of Color in Scaffolded Sensemaking around Algorithmic Fairness. In Proceedings of the 2023 ACM Conference on International Computing Education Research - Vol...

  38. [47]

    Antti Salovaara and Leevi Vahvelainen. 2025. Triangulating on Possible Futures: Conducting User Studies on Several Futures Instead of Only One. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New Y...

  39. [48]

    Christian Sandvig, Kevin Hamilton, Karrie Karahalios, and Cedric Langbort. 2014. Auditing algorithms: Research methods for detecting discrimination on internet platforms. Data and discrimination: converting critical concerns into productive inquiry 22, 2014 (2014), 4349–4357

  40. [49]

    Schafer, Kate Starbird, and Daniela K

    Joseph S. Schafer, Kate Starbird, and Daniela K. Rosner. 2023. Participatory Design and Power in Misinformation, Disinformation, and Online Hate Research. In Proceedings of the 2023 ACM Designing Interactive Systems Conference (Pittsburgh, PA, USA) (DIS ’23). Association for C...

  41. [50]

    Donald A Schön. 2017. The reflective practitioner: How professionals think in action . Routledge, New York

  42. [51]

    Hong Shen, Alicia DeVrio, Motahhare Eslami, and Kenneth Holstein. 2021. Everyday Algorithm Auditing: Understand- ing the Power of Everyday Users in Surfacing Harmful Algorithmic Behaviors. Proc. ACM Hum.-Comput. Interact. 5, CSCW2, Article 433 (Oct. 2021), 29 pages. doi:10.114...

  43. [52]

    Hong Shen, Haojian Jin, Ángel Alexander Cabrera, Adam Perer, Haiyi Zhu, and Jason I. Hong. 2020. Designing Alter- native Representations of Confusion Matrices to Support Non-Expert Public Understanding of Algorithm Performance. Proc. ACM Hum.-Comput. Interact. 4, CSCW2, Articl...

  44. [53]

    Jaemarie Solyst, Ellia Yang, Shixian Xie, Jessica Hammer, Amy Ogan, and Motahhare Eslami. 2024. Children’s Overtrust and Shifting Perspectives of Generative AI. In Proceedings of the 18th International Conference of the Learning Sciences - ICLS 2024, R. Lindgren, T. I. Asino, ...

  45. [54]

    Jaemarie Solyst, Ellia Yang, Shixian Xie, Amy Ogan, Jessica Hammer, and Motahhare Eslami. 2023. The Potential of Diverse Youth as Stakeholders in Identifying and Mitigating Algorithmic Bias for a Future of Fairer AI. Proc. ACM Hum.-Comput. Interact. 7, CSCW2, Article 364 (Oct....

  46. [55]

    Luhang Sun, Mian Wei, Yibing Sun, Yoo Ji Suh, Liwei Shen, and Sijia Yang. 2023. Smiling Women Pitching Down: Auditing Representational and Presentational Gender Biases in Image Generative AI. arXiv:2305.10566 [cs] doi:10. 48550/arXiv.2305.10566

  47. [56]

    Harini Suresh, Rajiv Movva, Amelia Lee Dogan, Rahul Bhargava, Isadora Cruxen, Angeles Martinez Cuba, Guilia Taurino, Wonyoung So, and Catherine D’Ignazio. 2022. Towards Intersectional Feminist and Participatory ML: A Case Study in Supporting Feminicide Counterdata Collection. ...

  48. [57]

    Latanya Sweeney. 2013. Discrimination in online ad delivery. Commun. ACM 56, 5 (May 2013), 44–54. doi:10.1145/ 2447976.2447990

  49. [58]

    Ying Tang, Hadar Ziv, and Sameer Patil. 2025. Learning to Work From Home: How Novice and Experienced Software Professionals Compare Online and In-person Collaboration. Proc. ACM Hum.-Comput. Interact. 9, 1, Article GROUP20 (Jan. 2025), 38 pages. doi:10.1145/3701199

  50. [59]

    TikTok. 2024. Effect Guidelines. https://effecthouse.tiktok.com/learn/guides/general/guidelines/effect-guidelines

  51. [60]

    Lauren Vogelstein, Vedya Konda, Deborah Fields, Yasmin Kafai, Luis Morales-Navarro, and Danaé Metaxa. 2025. Rapid Testing, Duck Lips, and Tilted Cameras: Youth Everyday Algorithm Auditing Practices with Generative AI Filters. In Proceedings of the 19th International Conference...

  52. [61]

    Kexin Bella Yang, Tomohiro Nagashima, Junhui Yao, Joseph Jay Williams, Kenneth Holstein, and Vincent Aleven

  53. [62]

    Robert K. Yin. 2012. Case study methods. In APA Handbook of Research Methods in Psychology, Vol. 2. Research Designs: Quantitative, Qualitative, Neuropsychological, and Biological , H. Cooper, P. M. Camic, D. L. Long, A. T. Panter, D. Rindskopf, and K. J. Sher (Eds.). American...

  54. [63]

    Robert K Yin. 2018. Case study research and applications

  55. [64]

    Manuel Zacklad. 2003. Communities of action: a cognitive and social approach to the design of CSCW systems. In Proceedings of the 2003 ACM International Conference on Supporting Group Work (Sanibel Island, Florida, USA) (GROUP ’03). Association for Computing Machinery, New Yor...

  56. [65]

    Yanzhe Zhang, Lu Jiang, Greg Turk, and Diyi Yang. 2023. Auditing Gender Presentation Differences in Text-to-Image Models. arXiv:2302.03675 [cs] doi:10.48550/arXiv.2302.03675

  57. [66]

    Yuhang Zhao. 2020. Analysis of TikTok’s Success Based on Its Algorithm Mechanism. In 2020 International Conference on Big Data and Social Sciences (ICBDSS) (2020-08). IEEE, Xi’an, China, 19–23. doi:10.1109/ICBDSS51270.2020.00012 Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7,...

  58. [2021]

    Can Crowds Customize Instructional Materials with Minimal Expert Guidance? Exploring Teacher-guided Crowdsourcing for Improving Hints in an AI-based Tutor. Proc. ACM Hum.-Comput. Interact. 5, CSCW1, Article 119 (April 2021), 24 pages. doi:10.1145/3449193

  59. [2022]

    End-User Audits: A System Empowering Communities to Lead Large-Scale Investigations of Harmful Algorithmic Behavior. Proc. ACM Hum.-Comput. Interact. 6, CSCW2, Article 512 (Nov. 2022), 34 pages. doi:10.1145/3555625

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.