Pith. sign in

REVIEW 4 major objections 5 minor 46 references

COGENT: A Curriculum-oriented Framework for Generating Grade-appropriate Educational Content

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read COGENT shows that adding curriculum structure to LLM prompts yields grade-appropriate science passages comparable to human-written ones.

desk verdict A genuinely useful pipeline for curriculum-aligned, readability-controlled content generation, but the 'grade-appropriate' claim rests on proxies, not on evidence from the target students. read the letter →

arxiv 2506.09367 v1 pith:FE6W6MUB submitted 2025-06-11 cs.CL cs.AI

classification cs.CLcs.AI
keywords curriculumalignmentgrade-appropriatereadabilityeducationalcontentgenerationlargelanguagemodelsscienceeducationwonder-basedlearningNGSSLLM-as-a-judge
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Generative language models can write fluent science passages, but without guidance they drift from curriculum standards and read at the wrong grade level. This paper argues that a structured prompt framework called COGENT can address both problems by feeding the model three curriculum elements (science concept, core idea, learning outcome), explicit readability targets (word count, vocabulary, sentence complexity), and a wonder-based question that frames the topic. Across three LLMs, COGENT produces passages rated significantly higher on curriculum alignment than both baseline prompts and human-written textbook passages, while matching target reading levels more closely in lower grades. The practical stake is that curriculum-aligned, grade-appropriate reading materials could be generated at scale, supplementing human authors and adapting to evolving standards.

What carries the argument

COGENT (Curriculum-Oriented Generation for Educational Content) is a prompt framework built from three components: curriculum formulation, controllable content generation, and multi-dimensional evaluation. The load-bearing mechanism is the hierarchical decomposition of NGSS standards into science concepts, core ideas, and grade-specific learning objectives, combined with explicit readability constraints (word count set to grade level times one hundred, a Flesch-Kincaid grade target, and vocabulary and sentence-complexity limits) and a 'wonder-based' inquiry question that turns each core idea into a curiosity trigger. This structure changes generation from open storytelling into standards-aligned explanation. The evaluation harness—LLM-as-a-judge scoring, expert teacher surveys, and statistical readability formulas—is what the paper uses to demonstrate that the mechanism works.

What would settle it

A classroom study in which students in grades 1-5 read COGENT, baseline, and human-written passages matched on topic and length, and then answer comprehension questions, would settle the claim: if COGENT passages do not yield comprehension at or above the human-written level in the target grades, the framework's central benefit is not confirmed.

Watch

Extended reading notes

Core claim

COGENT's central claim is that grade-appropriateness is not a property of the model alone but of the prompt's curriculum scaffolding. When the same LLM is given the science concept, core idea, and learning outcome plus explicit readability constraints, its output scores significantly higher on curriculum alignment than baseline prompts and human-written passages, without lowering comprehensibility. Readability metrics show COGENT passages stay closer to the intended grade level, especially in grades 1 and 2, while baseline prompts overshoot the target by roughly 2.5 grade levels. Expert elementary teachers rated COGENT passages at least as comprehensible as human-written passages. The paper reads this as evidence that LLM systems, with proper scaffolding, can serve as a complement to human expertise in educational content development.

Load-bearing premise

The load-bearing premise is that grade-appropriateness can be measured by readability formulas and expert or LLM ratings; because the study did not include elementary students, the claim would weaken if those proxies do not track real student comprehension.

Editorial extensions

If this is right

  • With COGENT-style scaffolding, smaller open models can produce curriculum-aligned passages, suggesting the approach does not require frontier models.
  • Adding curriculum guidance improves alignment without lowering comprehensibility, so standards and readability can be pursued together.
  • COGENT passages matched target reading levels more closely than baseline in lower grades, where readability mismatches are most harmful.
  • The framework is described as general, so the same decomposition could be adapted to other curricula and subjects beyond NGSS science.
  • Expert ratings support using COGENT-generated passages as a supplement to human-authored materials in elementary science education.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's limitations section acknowledges that elementary students did not evaluate the passages; whether readability formulas and expert ratings predict real student comprehension remains an open question the paper does not settle.
  • The wonder-based design aims at engagement, but engagement is only measured through expert judgment, not student behavior; a direct classroom comparison of curiosity or motivation would test that component.
  • If the curriculum decomposition were automated, the COGENT approach could be applied at scale to non-NGSS frameworks, turning any set of standards into readable passages without human curriculum mapping.
  • The paper observed lower curriculum-alignment ratings for all passage types in grades 3-5, which suggests that wonder-topic generation for higher-grade core ideas may need better matching to lift alignment further.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes COGENT, a curriculum-oriented framework for generating grade-appropriate elementary science reading passages. The framework decomposes NGSS standards into science concepts, core ideas, and learning objectives; conditions generation on these curriculum items along with word-count and Flesch-Kincaid readability targets; and adopts a 'wonder-based' question approach intended to enhance engagement. The evaluation compares three conditions (BASE, COGENT, and human-written passages) across three LLMs (Gemma-2-9B, GPT-4o, Claude-3.5-Sonnet) using LLM-as-a-judge (Claude-3.5-Sonnet), six elementary science teachers, and four readability formulas. The main reported results are that COGENT yields significantly higher curriculum-alignment scores than BASE and human references, comparable or higher comprehensibility, and readability scores closer to intended grade levels, especially in lower grades.

Significance. If the findings hold, COGENT offers a practical, standards-grounded approach for scaling curriculum-aligned reading materials, with strengths that include multi-model evaluation, independent teacher judgments, a clear hierarchical representation of curriculum knowledge, and an explicit engagement mechanism. The paper also makes a useful distinction between surface readability and comprehension. However, the strongest claims—'grade-appropriate' and 'comparable or superior to human references'—rest on evaluation choices that need scrutiny: the human baseline is confounded by topic selection, the automated judge is from the same model family as one generator, and no target-age students were tested. These gaps are fixable with additional analysis or revised claims, but they are load-bearing for the abstract's headline assertion.

major comments (4)
  1. [§5.2] The comparison against human-written passages is confounded by the choice of wonder topics. The paper states that 'the wonder topics extracted from the human references are not well-matched in these higher grades,' which explains why human passages receive lower curriculum-alignment ratings in grades 3–5. Since the human passages are judged against curriculum items that their own topics do not match, the human baseline is at an unfair disadvantage. To support the claim that COGENT is 'comparable or superior to human references,' the authors must either select human passages that are aligned with the same curriculum items (e.g., by using passages previously vetted against NGSS) or adopt an evaluation protocol that is neutral to topic origin (e.g., blind rating by independent judges). As it stands, the headline comparison is not a level playing field.
  2. [§4.2] The LLM-as-a-judge evaluation uses Claude-3.5-Sonnet as the scorer, which is also one of the three generators. This creates a potential same-family bias: the judge may systematically prefer outputs from its own model family due to shared style or training characteristics. The paper's claim that 'in our preliminary testing, Claude-3.5-Sonnet performs well as a consistent and accurate evaluator' is not backed by quantitative evidence. The authors should either use a judge from a different model family (e.g., a cross-evaluation design where each model is judged by a different-anthropic or different-openai model) or report inter-annotator agreement between the LLM judge and the six teachers on the same passages. Without such validation, Tables 3 and 4's automated scores are not fully convincing.
  3. [Limitations and §5.1] The central construct 'grade-appropriate' is not directly validated. The Limitations section acknowledges that 'we did not include elementary students in our sample analysis,' so grade-appropriateness is inferred from readability formulas and expert ratings. Because the COGENT prompt explicitly instructs the model to meet a Flesch-Kincaid grade level (Table 6), the closeness of COGENT's readability scores to intended grades in Figures 4 and 6 is partly an instruction-following result and does not by itself establish comprehension by target-grade readers. The paper itself notes the distinction between readability and comprehension (Section 5.1), yet the abstract claims 'grade-appropriate passages' without qualification. The authors should either temper the claim to 'curriculum-aligned and readability-controlled' or provide additional evidence linking readability to actual student comprehension, such as a small pilot study using comprehension questions or cloze tests with target-age children.
  4. [Tables 3–4, §5.1–5.2] The statistical reporting does not adequately support the significance claims. The unit of analysis is ambiguous: the grouped-generation results (Table 3) appear to treat each generated passage as independent, but multiple passages are generated per curriculum item (three samples per topic), so the independence assumption is violated. Additionally, the Bonferroni correction is not consistently applied: with three pairwise comparisons in Table 4, the corrected significance threshold is α = 0.0167, yet p-values such as .022, .029, .033, and .053 are reported with asterisks indicating significance. The authors should use mixed-effects models or cluster-robust tests that account for the nested structure, and should clearly state the Bonferroni correction procedure and provide effect sizes and confidence intervals.
minor comments (5)
  1. [§3.1, footnote 2] The readability and word-count targets are presented as design choices (word count = grade × 100), but Table 1 shows that human-written passages do not exactly follow this formula (e.g., grade 3 average is 319 words, not 300). Please clarify whether the targets are derived from human data or set heuristically, and discuss the potential consequence of this mismatch.
  2. [§5.2] The teacher evaluation (Section 5.3) covers only 15 passages (five grades × three conditions) and only one generation model (GPT-4o). The authors should justify this sample size and note that the expert results may not generalize to the other two models.
  3. [§5.1] The paper mentions that 'tested LLMs perform well (<6% averaged error rate) regarding Factual Correctness,' but this result is not shown in any figure or table. Please provide the supporting data or a citation to the evaluation protocol.
  4. [§5.2] The sentence 'the wonder topics extracted from the human references are not well-matched in these higher grades' is stated without explanation. Please elaborate on why the topics are mismatched and how this affects the interpretation of the human-baseline results.
  5. [Figure 4/6] The readability figures are printed at a small size and the legend colors are not described in the text. Consider adding a table with numerical values for readability metrics by grade and condition, as this would make the comparisons more transparent.

Circularity Check

1 steps flagged · score 6.0 of 10

The headline grade-appropriateness result is partly built into the COGENT prompt: the readability target is an input, then reported as evidence of grade-appropriateness; independent teacher ratings keep the claim from being fully circular.

  1. fitted input called prediction [Section 3.1 (readability control), Table 6 (COGENT prompt), results in Section 5.2 and Figure 6]
    "We thus indicate the word number and target readability level ... along with the curriculum input to ensure generated content matches students' reading abilities at each grade. ... Generate a 100-word reading passage around the Wonder Topic to teach students the Science Concept and Core Ideas, to meet the Learning Outcomes. Mix science and everyday language. - The generated text should meet the Flesch Kincaid Grade Level for elementary grade 1 students. ... COGENT produces passages closer to the intended grade level, while BASE generates passages largely above intended grade levels."

    The outcome used to substantiate the grade-appropriateness claim (Flesch-Kincaid grade level and word count) is exactly what the generation prompt instructs the model to produce. Section 3.1 and the Table 6 caption state that the word count is set as grade level times 100 and that Flesch Kincaid Grade Level is used for readability control; the COGENT prompt then says 'The generated text should meet the Flesch Kincaid Grade Level.' Section 5.2 cites COGENT's closer match to 'the intended grade level' as evidence of grade-appropriateness. That match is an instruction-following check, not an independent prediction. The teacher comprehensibility ratings are independent and support the overall claim, so the circularity is partial rather than total.

full rationale

The paper's central derivation chain is: inject curriculum items and readability targets into the prompt, then measure alignment to those same items and targets. The curriculum-alignment component is not a damaging circularity: COGENT is an intervention designed to inject curriculum information, so scoring how well the output reflects that information is a direct test of the intervention. The load-bearing circularity is in the grade-appropriate claim. Grade-appropriateness is operationalized by readability formulas, and the COGENT prompt explicitly requires the Flesch-Kincaid grade level and a word count derived from the grade level; the paper even calibrates the word count from human-written passages ('we set the word count to be the grade level multiplied by 100'). Reporting that COGENT lands near the target grade level is therefore partly a restatement of the prompt constraint. The Limitations section's admission that 'we did not include elementary students in our sample analysis' compounds this validity gap, since one of the proxy metrics (FK grade level) is both an input and the headline outcome. However, the paper also has genuinely independent support: six experienced elementary science teachers rated COGENT passages highly on comprehensibility and curriculum alignment, and COGENT was compared against human-written passages on the same instruments. No load-bearing self-citation or imported uniqueness theorem is present; the authors' self-citations in Related Work are background only. Overall, one central prediction reduces by construction while the rest of the evaluation retains independent content, so the circularity score is 6.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The evaluation rests on assumptions about the validity of NGSS mapping, readability formulas, LLM judge scores, and expert ratings. The paper releases no artifacts to independently verify these assumptions, and the word count and readability targets are hand-set parameters that influence all generated outputs.

free parameters (2)
  • word_count_target = grade level x 100 words (100 to 500)
    Set by hand from statistics of 50 human-written passages (Table 1) and used in all generation prompts; this is a chosen parameter, not learned from data.
  • readability_target = target grade as Flesch Kincaid Grade Level
    Prompts instruct the LLM to match the target grade level; derived from the authors' observation of human passages but selected manually.
assumptions (4)
  • domain assumption NGSS decomposition into science concepts, core ideas, and learning outcomes is a valid representation of curriculum standards.
    Section 3.1 uses NGSS as ground truth for generation and evaluation.
  • domain assumption Flesch-Kincaid, Gunning Fog, ARI, and Coleman-Liau formulas measure grade-appropriate readability for elementary students.
    Used as 'Text Readability' metrics in Section 3.2 and Figures 4 and 6.
  • domain assumption Claude-3.5-Sonnet provides valid LLM-as-judge scores for curriculum alignment and comprehensibility.
    Section 4.2 states it 'performs well' in preliminary testing, but no calibration or inter-rater agreement is reported.
  • domain assumption Teacher ratings on 15 passages are representative of pedagogical quality.
    Expert analysis in Section 5.3 uses 3 passages per grade; no inter-rater reliability is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of COGENT: A Curriculum-oriented Framework for Generating Grade-appropriate Educational Content." pith.science (2026). https://pith.science/paper/FE6W6MUB

@misc{pith2026250609367,
  author       = {Pith},
  title        = {Pith review of: COGENT: A Curriculum-oriented Framework for Generating Grade-appropriate Educational Content},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FE6W6MUB}},
  note         = {Machine review of arXiv:2506.09367}
}
read the original abstract

While Generative AI has demonstrated strong potential and versatility in content generation, its application to educational contexts presents several challenges. Models often fail to align with curriculum standards and maintain grade-appropriate reading levels consistently. Furthermore, STEM education poses additional challenges in balancing scientific explanations with everyday language when introducing complex and abstract ideas and phenomena to younger students. In this work, we propose COGENT, a curriculum-oriented framework for generating grade-appropriate educational content. We incorporate three curriculum components (science concepts, core ideas, and learning objectives), control readability through length, vocabulary, and sentence complexity, and adopt a ``wonder-based'' approach to increase student engagement and interest. We conduct a multi-dimensional evaluation via both LLM-as-a-judge and human expert analysis. Experimental results show that COGENT consistently produces grade-appropriate passages that are comparable or superior to human references. Our work establishes a viable approach for scaling adaptive and high-quality learning resources.

Figures

Figures reproduced from arXiv: 2506.09367 by the authors.

Figure 1
Figure 1. Overview of the framework of curriculum-oriented generation for educational content (COGENT). [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Our curriculum decomposition example grounded in the Next Generation Science Standards (NGSS), [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Curriculum alignment scores (left) and comprehensibility scores (right) of Gemma-2-9B, GPT-4o, and [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Results on four readability metrics of LLM-generated passages using BASE and COGENT framework. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Results on curriculum alignment and comprehensibility of Human, BASE, and COGENT. [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Results on readability metrics of human-written passages, BASE, and COGENT framework. [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Expert analysis: curriculum alignment comparison of Human, BASE, and COGENT. [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Expert analysis: comprehensibility score comparison among Human, BASE, and COGENT. [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

46 extracted references · 37 canonical work pages

  1. [1]

    Fouad Abd-El-Khalick, Saouma Boujaoude, Richard Duschl, Norman G Lederman, Rachel Mamlok-Naaman, Avi Hofstein, Mansoor Niaz, David Treagust, and Hsiao-lin Tuan. 2004. https://doi.org/10.1002/sce.10118 Inquiry in science education: International perspectives . Science Education, 88:397--419

  2. [2]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774

  3. [3]

    Lorin W Anderson. 2002. Curriculum alignment: A re-examination. Theory into Practice, 41(4):255--260

  4. [4]

    Apler J Bansiong. 2019. https://doi.org/10.1080/2331186X.2019.1706395 Readability, content, and mechanical feature analysis of selected commercial science textbooks intended for third grade filipino learners . Cogent Education, 6(1):1706395

  5. [5]

    Isabel L Beck, Margaret G McKeown, Gale M Sinatra, and Jane A Loxterman. 1991. https://doi.org/10.2307/747763 Revising social studies text from a text-processing perspective: Evidence of improved comprehensibility . Reading Research Quarterly, 26(3):251--276

  6. [6]

    Adele Berndt and Jane P. Wayland. 2014. Evaluating the readability of marketing research textbooks: an international comparison. Journal of International Education in Business, 7(1):47--59

  7. [7]

    Eric J Blown and Tom GK Bryce. 2017. Switching between everyday and scientific language. Research in Science Education, 47:621--653

  8. [8]

    Rodger W Bybee. 2014. Ngss and the next generation of science teachers. Journal of science teacher education, 25(2):211--221

Show all 46 references
  1. [9]

    Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020. Evaluation of text generation: A survey. arXiv:2006.14799

  2. [10]

    Christine Chin and David E Brown. 2002. Student-generated questions: A meaningful aspect of learning in science. International Journal of Science Education, 24(5):521--549

  3. [11]

    Department for Education . 2014. National curriculum. https://www.gov.uk/government/collections/national-curriculum. The national curriculum for England to be taught in all local-authority-maintained schools. Introduced September 2014, with English and maths coming into force ...

  4. [12]

    John Dewey. 1986. Experience and education. In The Educational Forum, pages 241--252. Taylor & Francis Group

  5. [13]

    Rudolph Flesch. 1948. A new readability yardstick. Journal of Applied Psychology, 32(3):221

  6. [14]

    Andrew Gilbert and Christie C Byers. 2017. Wonder as a tool to engage preservice elementary teachers in science learning and teaching. Science Education, 101(6):907--928

  7. [15]

    Robert Gunning. 1968. The Technique of Clear Writing, 2nd edition. McGraw-Hill, New York

  8. [16]

    Simon Hughes and Minseok Bae. 2023. https://github.com/vectara/hallucination-leaderboard Vectara hallucination leaderboard

  9. [17]

    Jamie J Jirout. 2020. Supporting early scientific thinking through curiosity. Frontiers in Psychology, 11:1717

  10. [18]

    Md Rayhan Kabir and Fuhua Lin. 2023. An LLM -powered adaptive practicing system. In LLM@AIED, pages 43--52

  11. [19]

    George R Klare. 1974. Assessing readability. Reading Research Quarterly, pages 62--102

  12. [20]

    Bor-Chen Kuo, Frederic TY Chang, and Zong-En Bai. 2023. Leveraging LLMs for adaptive testing and learning in Taiwan adaptive learning platform ( TALP ). In LLM@AIED, pages 101--110

  13. [21]

    George Lakoff and Mark Johnson. 1980. The metaphorical structure of the human conceptual system. Cognitive Science, 4(2):195--208

  14. [22]

    Soohwan Lee and Ki-Sang Song. 2024. Teachers' and students' perceptions of AI -generated concept explanations: Implications for integrating generative AI in computer science education. Computers and Education: Artificial Intelligence, 7:100283

  15. [23]

    Unggi Lee, Haewon Jung, Younghoon Jeon, Younghoon Sohn, Wonhee Hwang, Jewoong Moon, and Hyeoncheol Kim. 2024. https://doi.org/10.1007/s10639-023-12249-8 Few-shot is enough: exploring ChatGPT prompt engineering method for automatic question generation in english education . Edu...

  16. [24]

    Minzhi Li, Zhengyuan Liu, Shumin Deng, Shafiq Joty, Nancy Chen, and Min-Yen Kan. 2025. https://aclanthology.org/2025.coling-main.156/ D n A -eval: Enhancing large language model evaluation through decomposition and aggregation . In Proceedings of the 31st International Confere...

  17. [25]

    Ta Lin Liau, Carolyn B Bassin, Clessen J Martin, and Edmund B Coleman. 1976. Modification of the coleman readability formulas. Journal of Reading Behavior, 8(4):381--386

  18. [26]

    Markus Lindholm. 2018. Promoting curiosity? possibilities and pitfalls in science education. Science & Education, 27:987--1002

  19. [27]

    Zhengyuan Liu, Stella Xin Yin, and Nancy Chen. 2024 a . https://doi.org/10.18653/v1/2024.sigdial-1.43 Optimizing code-switching in conversational tutoring systems: A pedagogical framework and evaluation . In Proceedings of the 25th Annual Meeting of the Special Interest Group ...

  20. [28]

    Zhengyuan Liu, Stella Xin Yin, Carolyn Lee, and Nancy F Chen. 2024 b . Scaffolding language learning via multi-modal tutoring systems with pedagogical instructions. In 2024 IEEE Conference on Artificial Intelligence (CAI), pages 1258--1265. IEEE

  21. [29]

    Zhengyuan Liu, Stella Xin Yin, Geyu Lin, and Nancy F. Chen. 2024 c . https://doi.org/10.18653/v1/2024.emnlp-main.37 Personality-aware student simulation for conversational intelligent tutoring systems . In Proceedings of the 2024 Conference on Empirical Methods in Natural Lang...

  22. [30]

    Ministry of Education Singapore . 2023. Primary school curriculum and subjects. https://www.moe.gov.sg/primary/curriculum. Last updated: 02 Mar 2023. The primary school curriculum is designed to give children of school-going age a strong foundation in learning

  23. [31]

    National Research Council . 2000. Inquiry and the national science standards. National Academy Press, Washington, DC

  24. [32]

    Mark Sadoski, Ernest T Goetz, and Maximo Rodriguez. 2000. Engaging texts: Effects of concreteness on comprehensibility, interest, and recall in four text types. Journal of Educational Psychology, 92(1):85

  25. [33]

    Swarnadeep Saha, Omer Levy, Asli Celikyilmaz, Mohit Bansal, Jason Weston, and Xian Li. 2024. Branch-solve-merge improves large language model evaluation and generation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Lin...

  26. [34]

    Edgar A Smith and RJ Senter. 1967. Automated readability index. Technical Report Vol. 66, No. 220, Aerospace Medical Research Laboratories, Aerospace Medical Division, Air Force Systems Command

  27. [35]

    David Squires. 2012. Curriculum alignment research suggests that alignment can improve student achievement. The Clearing House: A Journal of Educational Strategies, Issues and Ideas, 85(4):129--135

  28. [36]

    NGSS Lead States. 2013. Next generation science standards: For states, by states. National Academies Press

  29. [37]

    Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, L \'e onard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ram \'e , et al. 2024. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118

  30. [38]

    Shen Wang, Tianlong Xu, Hang Li, Chaoli Zhang, Joleen Liang, Jiliang Tang, Philip S Yu, and Qingsong Wen. 2024. Large language models for education: A survey and outlook. arXiv preprint arXiv:2403.18105

  31. [39]

    Sheida White and John Clement. 2001. Assessing the lexile framework: Results of a panel meeting. National Center for Education Statistics

  32. [40]

    Leoniek Wijngaards-de Meij and Sigrid Merx. 2018. Improving curriculum alignment and achieving learning goals by making the curriculum visible. International Journal for Academic Development, 23(3):219--231

  33. [41]

    Changrong Xiao, Sean Xin Xu, Kunpeng Zhang, Yufang Wang, and Lei Xia. 2023. Evaluating reading comprehension exercises generated by LLMs : A showcase of ChatGPT in education applications. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational App...

  34. [42]

    Lixiang Yan, Lele Sha, Linxuan Zhao, Yuheng Li, Roberto Martinez-Maldonado, Guanliang Chen, Xinyu Li, Yueqiao Jin, and Dragan Ga s evi \'c . 2024. https://doi.org/10.1111/bjet.13370 Practical and ethical challenges of large language models in education: A systematic scoping re...

  35. [43]

    Mostafa Zamanian and Pooneh Heydari. 2012. Readability of texts: State of the art. Theory & Practice in Language Studies (TPLS), 2(1)

  36. [44]

    Eric Zelikman, Wanjing Ma, Jasmine Tran, Diyi Yang, Jason Yeatman, and Nick Haber. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.135 Generating and evaluating tests for k-12 students with language model simulations: A case study on sentence reading efficiency . In Proceedi...

  37. [45]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before...

  38. [46]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.