REVIEW 3 major objections 5 minor 33 references
A single accessibility-configured prompt raised the mean WCAG compliance of AI-generated teaching materials from 24.2% to 96.7% across five content types in a first exploratory run.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-05 00:21 UTC pith:WCPKPPSF
load-bearing objection A genuinely reusable protocol for WCAG-based evaluation of AI-generated educational content; the pilot numbers are directionally suggestive but carry no statistical weight. the 3 major comments →
A Protocol for Evaluating the Accessibility of AI-Generated Educational Materials: Prompt Configuration, WCAG-Derived Criteria, and Content Overload
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim, on the paper's own terms, is that a generative AI model reproduces inaccessible practices by default but can meet most operationalized WCAG Level AA criteria when the prompt explicitly states them; the same model, same task, and one added configuration step moved the mean compliance score from 24.2% to 96.7% across the five content types in the exploratory run. The protocol is the main contribution: a content-type-specific WCAG rubric, a prompt bank with generic, configured, and persistent-profile conditions, and an expert-panel heuristic validation procedure. The paper also identifies two barriers that conventional WCAG checklists miss: information overload in dense AI-sy
What carries the argument
The load-bearing object is the evaluation protocol: for each content type, a WCAG 2.2 Level AA-derived rubric with criteria for documents (headings, table headers, reading order, alt text, contrast, bookmarked TOC, form labels), slides (placeholder reading order, alt text, keyboard operability, contrast, flashing limits, unique titles), images/infographics (text alternative, embedded-text contrast, density, color reliance), audio (transcript, pacing, no visual-only references), and video (captions, audio description, transcript, lip-sync/viseme alignment). Three experimental conditions are compared on the same tool: generic prompt, WCAG-configured prompt, and a persistent accessibility profi
Load-bearing premise
The reported gain rests on a single non-blind evaluator's rubric scoring of one generated artifact per condition and content type, plus local open-source tool conversions standing in for the specialized commercial platforms; if those measurements are noisy or biased, the 24.2% to 96.7% gap could be an artifact.
What would settle it
Run the full protocol with at least three blind evaluators, at least ten artifacts per condition per content type, and the actual commercial tools named in the paper; the central claim would be falsified if the WCAG-configured condition no longer clearly outperforms the generic condition, or if compliance scores on the commercial platforms differ drastically from the open-source proxy scores.
If this is right
- If the central claim holds, a one-sentence accessibility requirement in the prompt can substitute for a separate post-generation remediation step for documents, slides, images/infographics, and audio.
- Video is the content type where configuration helped least in this pilot; the lip-sync barrier suggests that some video accessibility defects are not reachable by prompt text alone.
- Accessibility is a property of each generated content item, not just of the platform that hosts it; an LMS can be compliant while the content uploaded into it is not.
- A persistent accessibility profile, if it reduces variance across requests, would address the practical failure mode where busy creators shorten or forget the full prompt.
- The open protocol and rubric are designed to be reused: other teams can apply the same criteria to their own tools and report comparable numbers.
Where Pith is reading between the lines
- The 24.2% to 96.7% numbers come from a single non-blind evaluator, one artifact per condition, and local open-source proxies; I would treat them as existence evidence that configuration changes output, not as a precise effect-size estimate for commercial tools.
- A natural extension is to validate the information-density indicator against user performance or self-reported cognitive load, connecting RQ4 to broader cognitive-load research rather than relying on expert rating alone.
- The persistent-profile hypothesis could be sharpened by comparing the variance of compliance scores across repeated varied requests, not just the mean: if profiles lower variance, their value is consistency; if only the mean rises, they are mainly a convenience.
- Provider accessibility claims could be tested with the same rubric: a specialized tool that markets an accessibility checker would face a harder test than a general-purpose assistant if both score low under a generic prompt.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a protocol for evaluating the accessibility of AI-generated educational materials across five content types (documents, slides, images/infographics, audio, video). The protocol defines three conditions: a generic prompt (A), a single WCAG-configured prompt (B), and a persistent accessibility profile (C). Evaluation combines an operationalized WCAG-derived rubric with expert heuristic validation. The main contribution is methodological: a detailed, openly released evaluation instrument, prompt bank, PRISMA-informed review protocol, and a concrete recipe for implementing Condition C. A first exploratory application using one model, one artifact per condition and content type, and single-evaluator scoring reports an increase in rubric compliance from an unweighted mean of 24.2% under Condition A to 96.7% under Condition B; Condition C was not exercised. The paper explicitly acknowledges the exploratory scope and limitations in Sections 5.6 and 6.
Significance. If the protocol is adopted, it would fill a real gap: existing studies evaluate accessibility of AI-generated content for isolated content types, while this paper provides a unified, content-type-specific instrument with operationalized WCAG criteria and an explicit overload dimension. The appendices are unusually complete: full prompt bank, scoring rubric with rating guidance, PRISMA-style search strings, expert-validation instructions, and an implementation recipe for a persistent profile. The honest and detailed limitation statements are a strength. The reported 24.2% to 96.7% improvement, however, is only an illustrative pilot result. Because it is based on one non-blind evaluator, one artifact per condition, and local proxies for commercial tools, the quantitative magnitude carries no statistical weight. The protocol’s value does not depend on this particular number, but the abstract currently presents the number as a headline finding, which overstates the evidence.
major comments (3)
- [Abstract; §4.1; Table 7; §5.6] The headline quantitative claim, a rise from 24.2% to 96.7% compliance, rests on one artifact per condition and content type scored by a single non-blind evaluator with spot-checks, not the blinded multi-rater expert panel specified in §3.4. The paper acknowledges this in §5.6 and §6, but the Abstract still presents the numbers as the result. For a protocol whose contribution is reproducibility, the pilot measurement pipeline needs either to be reported as purely illustrative in the Abstract as well as the body, or to be re-run with repeated artifacts, blind scoring, and inter-rater agreement. Without that, the magnitude is not reliable.
- [§3.6; §4.1] Four of the five content types were evaluated after conversion using local open-source proxies (python-docx, python-pptx, Pillow, eSpeak, FFmpeg) rather than the specialized commercial tools named in §3.2. Any accessibility structure introduced by the conversion process, for example heading tags in python-docx or metadata in FFmpeg, is attributed to the AI tool. This means the exploratory numbers do not validate the tools listed in Table 3. The paper’s target design is sound, but the pilot results should be explicitly restricted to the proxy toolchain, and confirmation on the named platforms remains future work.
- [§3.4; §3.6; Appendix B; Appendix C] The pilot did not implement the expert heuristic validation panel described in §3.4, including blinding and evaluators with lived experience of disability. Because the Condition B prompts in Appendix B closely mirror the rubric criteria in Appendix C, high Condition B scores may partly reflect prompt adherence rather than genuine accessibility for users with disabilities. This is not a circularity in the rubric itself, but it is a construct-validity concern for the exploratory comparison. The protocol already specifies the remedy; the pilot should either adopt it or the quantitative comparison should be labeled as a demonstration of prompt-following only.
minor comments (5)
- [Abstract; Table 7] The Abstract calls the 24.2% and 96.7% values a 'pooled mean,' but Table 7 reports an unweighted mean across five content types. 'Pooled mean' normally implies weighting by the number of criteria, which would give different values. The terminology should be aligned with the actual calculation.
- [§4.1] The sentence 'one artifact was evaluated per condition and content type, not a repeated or averaged sample' is good, but it appears only after the results table. Consider placing this caveat immediately before the reported percentages, and repeating it in the Abstract if the numbers are retained.
- [Declarations] The pilot data-recording spreadsheet and generated outputs are not released; only the protocol, prompt bank, and rubric are open. For full reproducibility of the reported numbers, the raw scoring sheets or at least the unweighted per-criterion scores for each artifact should be made available, or the claim of reproducibility should be scoped to the instrument rather than the pilot results.
- [Appendix A.6] The PRISMA flow diagram numbers for identification and screening are marked [n] and require a formal database re-run. This is honest, but the caption or surrounding text should state explicitly that the current review is a narrative review formalized after the fact, so readers do not mistake the protocol for an executed systematic review.
- [Table 5] The video Condition B prompt requires captions 'synchronized within 100ms.' WCAG 2.2 does not specify a 100 ms tolerance; if this is intended as an operationalization, it should be justified or softened to 'synchronized' with the 100 ms value reported as a study choice.
Circularity Check
No significant circularity: the compliance score is a directly measured rubric percentage (App. C.2), and the paper's only load-bearing premises rest on external evidence, not self-citations.
full rationale
The paper's load-bearing empirical result—a compliance score rising from 24.2% to 96.7%—is a direct measurement, not a derivation. Table 7 scores are computed from the rubric in Appendix C via the explicit formula in C.2; no parameter is fitted, no hidden quantity is renamed as a prediction, and no equation reduces to another equation by construction. Condition B's high scores reflect the fact that the configured prompt enumerates the same operationalized WCAG criteria that the rubric later checks; this is an intentional intervention-fidelity design, but it is not circular in the logical sense because the rubric is independently anchored to WCAG and the pilot is explicitly reported as an exploratory measurement rather than a derivation. The paper itself scopes the pilot in §3.6, §5.6, and §6: one model, one artifact per condition, a non-blind single evaluator, and local open-source proxy tooling. Those are validity and robustness concerns (expectation bias, sample size, ecological validity), not circularity. Self-citations (refs. 1–3, 5–9) appear only in background framing about e-learning and orchestrated content creation; the accessibility-default and configuration-help premises are supported by external studies ([10]–[19]) and by the paper's own independent pilot. No uniqueness theorem or ansatz is imported from the authors' prior work. The only equation in the paper is the definition of the compliance percentage (App. C.2), which is a stated measurement convention and cannot be circular. The abstract's 'pooled mean' label is corrected by the table note as an unweighted mean; this is a reporting imprecision, not a circular step. Overall score 1: no significant circularity; the minor self-citations present are not load-bearing.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption WCAG 2.2 Level AA is the appropriate normative target for judging accessibility of the evaluated content.
- domain assumption The operationalized criteria in Table 6 faithfully capture WCAG success criteria for each content type.
- ad hoc to paper The information-density indicator (1-5 scale in Appendix C.4) is a valid measure of the content-overload dimension (RQ4).
- domain assumption The Condition A prompts represent typical current usage of AI tools for educational content creation.
- domain assumption The local open-source conversions (python-docx, LibreOffice, etc.) faithfully reproduce the accessibility properties of the specialized commercial platforms' outputs.
Cite this review
Pith. "Pith review of A Protocol for Evaluating the Accessibility of AI-Generated Educational Materials: Prompt Configuration, WCAG-Derived Criteria, and Content Overload." pith.science (2026). https://pith.science/paper/WCPKPPSF
@misc{pith2026260800749,
author = {Pith},
title = {Pith review of: A Protocol for Evaluating the Accessibility of AI-Generated Educational Materials: Prompt Configuration, WCAG-Derived Criteria, and Content Overload},
year = {2026},
howpublished = {\url{https://pith.science/paper/WCPKPPSF}},
note = {Machine review of arXiv:2608.00749}
}
read the original abstract
Generative AI tools increasingly produce educational materials: documents, slides, images, audio, and video, yet little is known about whether this content meets accessibility requirements. This paper presents a protocol for evaluating the accessibility of AI-generated educational materials against the Web Content Accessibility Guidelines (WCAG), across five content types and multiple tools. The protocol compares three conditions applied to the same tool: a generic instruction with no accessibility language; a single prompt explicitly configured with WCAG criteria; and a persistent, reusable accessibility profile loaded once rather than re-specified each time. Evaluation combines a WCAG rubric per content type with heuristic validation by accessibility experts, addressing a known limitation of automated scanners. Prior evidence shows generative AI tools reproduce inaccessible practices by default, and that explicit configuration measurably improves compliance. This paper extends that discussion to a dimension WCAG checklists miss: visual and informational overload common in AI-synthesized content. The main contribution is methodological: a reproducible protocol and open evaluation instrument, with cases documenting the barriers non-configured AI content creates for people with disabilities. A first exploratory application is also reported: a rubric-based compliance score rose from a pooled mean of 24.2% under the generic condition to 96.7% under the WCAG-configured condition, using one model, one artifact per condition and type, and single-evaluator scoring; the persistent-profile condition was not exercised. The protocol is a reusable resource for the community to apply and extend at full benchmark scale. Beyond the protocol, this work aims to raise awareness of accessibility obstacles generative AI can introduce, and encourage creators to consider accessibility when using these tools.
Reference graph
Works this paper leans on
-
[1]
Morales-Chan, M., Amado-Salvatierra, H. R., & Hernández-Rizzardini, R. (2024). AI-driven content creation: Revolutionizing educational materials. InProceedings of the Eleventh ACM Conference on Learning @ Scale (L@S ’24)(pp. 556–558). ACM. https://doi.org/10.1145/3657604. 3664640
doi:10.1145/3657604 2024
-
[2]
Morales-Chan, M., Amado-Salvatierra, H. R., Medina, J. A., Barchino, R., Hernández-Rizzardini, R., & Teixeira, A. M. (2024). Personalized feedback in Massive Open Online Courses: Harnessing the power of LangChain and OpenAI API.Electronics, 13(10), 1960.https://doi.org/10.3390/ electronics13101960
work page 2024
-
[3]
Morales Chan, M., Amado-Salvatierra, H. R., & Hernández Rizzardini, R. (2023). Optimizing the design, pedagogical decision-making and development of MOOCs through the use of AI-based tools. In Post-Covid Prospects for Massive Open Online Courses: 8th European MOOCs Stakeholders Summit (EMOOCs 2023)(pp. 95–103). Universitätsverlag Potsdam
work page 2023
-
[4]
(2023).Web Content Accessibility Guidelines (WCAG) 2.2
World Wide Web Consortium (W3C). (2023).Web Content Accessibility Guidelines (WCAG) 2.2. https://www.w3.org/TR/WCAG22/
work page 2023
-
[5]
R., Hernández, R., & Hilera, J
Amado-Salvatierra, H. R., Hernández, R., & Hilera, J. R. (2012). Implementation of accessibility standards in the process of course design in virtual learning environments.Procedia Computer Science, 14, 363–370.https://doi.org/10.1016/j.procs.2012.10.041
-
[6]
Martín, J. L., Amado-Salvatierra, H. R., & Hilera González, J. R. (2016). MOOCs for all: Evaluating the accessibility of top MOOC platforms.International Journal of Engineering Education, 32(5-B), 2274–2283
work page 2016
-
[7]
R., Hernández, R., & Hilera, J
Amado-Salvatierra, H. R., Hernández, R., & Hilera, J. R. (2014). Teaching and promoting web acces- sibility in virtual learning environments: A staff training experience in Latin-America. In2014 IEEE Frontiers in Education Conference (FIE). https://doi.org/10.1109/FIE.2014.7044392
-
[8]
Amado-Salvatierra, H. R., & Hilera, J. R. (2015). Towards an approach for an accessible and inclusive virtual education using ESVI-AL project results.Interactive Technology and Smart Education, 12(3), 158–168.https://doi.org/10.1108/ITSE-04-2015-0005
-
[9]
H., Chang, V ., Gütl, C., & Amado-Salvatierra, H
Rizzardini, R. H., Chang, V ., Gütl, C., & Amado-Salvatierra, H. R. (2013). An open online course with accessibility features. InProceedings of EdMedia 2013 — World Conference on Educational Media and Technology(pp. 635–643). AACE
work page 2013
-
[10]
Aljedaani, W., Habib, A., Aljohani, A., Eler, M. M., & Feng, Y . (2024). Does ChatGPT generate accessible code? Investigating accessibility challenges in LLM-generated source code. InProceedings of the 21st International Web for All Conference (W4A ’24). ACM. https://doi.org/10.1145/ 3677846.3677854
-
[11]
Abu Doush, I., & Kassem, R. (2025). Can generative AI create accessible web code? A benchmark analysis of AI-generated HTML against accessibility standards.Universal Access in the Information Society, 24(4), 3483–3506.https://doi.org/10.1007/s10209-025-01250-2
-
[12]
Abu Doush, I., & Kassem, R. (2024). Evaluating AI-generated web code for accessibility compliance: A metric-driven approach. InProceedings of the 11th International Conference on Software Development and Technologies for Enhancing Accessibility and Fighting Info-exclusion (DSAI 2024). ACM.https: //doi.org/10.1145/3696593.3696614 34
-
[13]
WebAIM. (2025).The WebAIM Million: The 2025 report on the accessibility of the top 1,000,000 home pages.https://webaim.org/projects/million/2025
work page 2025
-
[14]
Alshaigy, B., & Grande, V . (2024). Forgotten again: Addressing accessibility challenges of generative AI tools for people with disabilities. InAdjunct Proceedings of the 2024 Nordic Conference on Human- Computer Interaction (NordiCHI Adjunct ’24). ACM. https://doi.org/10.1145/3677045. 3685493
-
[15]
Vera-Amaro, G., & Rojano-Cáceres, J. R. (2025). Accessible web content generation using LLMs: An empirical study on prompting strategies and template-guided remediation.IEEE Latin America Transactions, 23(12), 1230–1239
work page 2025
-
[16]
(2024).ACCESS: Prompt engineering for automated web accessibility violation corrections
Huang, C., Ma, A., Vyasamudri, S., Puype, E., Kamal, S., Belza Garcia, J., Cheema, S., & Lutz, M. (2024).ACCESS: Prompt engineering for automated web accessibility violation corrections. arXiv. https://doi.org/10.48550/arXiv.2401.16450
-
[17]
Shen, Y ., Zhang, H., Shen, Y ., Wang, L., Shi, C., Du, S., & Tao, Y . (2025). AltGen: AI-driven alt text generation for enhancing EPUB accessibility. InProceedings of the 2025 International Conference on Artificial Intelligence and Computational Intelligence (AICI 2025). ACM.https://doi.org/10. 1145/3730436.3730449
-
[18]
Schmitt-Koopmann, F. M., Huang, E. M., Hutter, H.-P., & Darvishy, A. (2025). Towards more accessible scientific PDFs for people with visual impairments: Step-by-step PDF remediation to improve tag accuracy. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM. https://doi.org/10.1145/3706598.3713084
arXiv 2025
-
[19]
Baglodi, D., Martincic, B., Sinclair, N., & Walker, B. N. (2025). Automated, context-aware alt text generation for educational documents using large language models. InProceedings of the 27th International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS ’25). ACM. https://doi.org/10.1145/3663547.3759748
- [20]
-
[21]
Rose, D. H., & Meyer, A. (2002).Teaching every student in the digital age: Universal Design for Learning. ASCD
work page 2002
-
[22]
Ahmed, A., Fresco, M., Forsberg, F., & Grotli, H. (2025).From code to compliance: Assessing ChatGPT’s utility in designing an accessible webpage — A case study(v2). arXiv. https://doi. org/10.48550/arXiv.2501.03572
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2501.03572 2025
-
[23]
Palmer, Z. B., & Oswal, S. K. (2025). Constructing websites with generative AI tools: The accessi- bility of their workflows and products for users with disabilities.Journal of Technical Writing and Communication.https://doi.org/10.1177/10506519241280644
- [24]
-
[25]
Kumar, A., Padath, T., & Wang, L. L. (2025). Benchmarking PDF accessibility evaluation: A dataset and framework for assessing automated and LLM-based approaches for accessibility testing. InProceedings of the 27th International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS ’25). ACM.https://doi.org/10.1145/3663547.3746380 35
arXiv 2025
-
[26]
(2026).LLM-driven accessible interface: A model- based approach
Jerry, B., Moreno, L., Francisco, V ., & Hervás, R. (2026).LLM-driven accessible interface: A model- based approach. arXiv.https://doi.org/10.48550/arXiv.2601.06616
-
[27]
Vásquez-Rodríguez, L., Cuenca-Jiménez, P.-M., Morales-Esquivel, S., & Alva-Manchego, F. (2022). A benchmark for neural readability assessment of texts in Spanish. InProceedings of the Workshop on Text Simplification, Accessibility, and Readability (TSAR-2022)(pp. 188–198). Association for Computational Linguistics.https://doi.org/10.18653/v1/2022.tsar-1.18
-
[28]
(2024).Exploring large language models to generate Easy to Read content
Martínez, P., Moreno, L., & Ramos, A. (2024).Exploring large language models to generate Easy to Read content. arXiv.https://doi.org/10.48550/arXiv.2407.20046
-
[29]
(2025).Text adaptation to plain language and Easy Read via automatic post-editing cycles
Calleja, J., Ponce, D., & Etchegoyhen, T. (2025).Text adaptation to plain language and Easy Read via automatic post-editing cycles. arXiv.https://doi.org/10.48550/arXiv.2509.11991
-
[30]
Malakul, S., & Park, I. (2023). The effects of using an auto-subtitle system in educational videos to facilitate learning for secondary school students: Learning comprehension, cognitive load, and satisfaction.Smart Learning Environments, 10, Article 3. https://doi.org/10.1186/ s40561-023-00224-2
work page 2023
-
[31]
Donnelly, M. (2026). Introducing the Edu-GenAI Rubric: A theory-informed tool for assessing the educational value of large language models and AI media generators.Education Sciences, 16(5), 706. https://doi.org/10.3390/educsci16050706
-
[32]
He, P., Li, Z., Wang, Z., Xiong, J., & Li, T. (2026).Judging the judges: Human validation of multi-LLM evaluation for high-quality K-12 science instructional materials. arXiv. https://doi.org/10. 48550/arXiv.2602.13243
-
[33]
Schmitt-Koopmann, F. M., Huang, E. M., & Darvishy, A. (2022). Accessible PDFs: Applying artificial intelligence for automated remediation of STEM PDFs. InProceedings of the 24th International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS ’22). ACM. https://doi.org/ 10.1145/3517428.3550407 36
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.