Pith. sign in

REVIEW 4 major objections 5 minor 53 references

Validating the Effectiveness of a Large Language Model-based Approach for Identifying Children's Development across Various Free Play Settings in Kindergarten

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A large language model can identify kindergarteners' developmental abilities from their own free-play narratives with over 90 percent accuracy in most domains.

desk verdict A useful pilot dataset and evaluation protocol, but the central claim of identifying development is unsupported: the performance score is just LLM label frequency, and the cross-setting statistics violate the independence assumption. read the letter →

arxiv 2505.03369 v1 pith:DIBBI3PM submitted 2025-05-06 cs.AI cs.CY

classification cs.AIcs.CY
keywords LargelanguagemodelsLearninganalyticsEarlychildhoodeducationChilddevelopmentFreeplaySelf-narrativesPerformancescoringsettings
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that a large language model can act as a reliable developmental assessor in kindergarten free play by reading children's own accounts of what they did. The authors collected 2,224 self-narratives from 29 children across four play areas, had the LLM tag eight abilities in four domains, and asked eight early-childhood professionals to judge a random sample of those tags. They report accuracy above 90 percent for semantic consistency, relevance, and combined accuracy in cognitive, motor, and social abilities, with lower performance in emotional abilities. They also report statistically significant differences across play settings for seven of eight abilities, concluding the approach is highly effective for identifying development in free play. If this holds, teachers could monitor developmental trajectories from children's own words instead of relying only on observation and memory.

What carries the argument

The machinery is a five-step pipeline: collect children's narratives, proofread and anonymize them, prompt the LLM to select abilities and describe observed behavior, format the output into structured ability-performance records, and compute performance scores. The load-bearing definition is the performance score, $$ ext{score} = rac{ ext{number of activity records in which the LLM infers that ability}}{ ext{total activity records for that child in the period}}.$$ The ability categories come from an ECDI2030-inspired framework covering numeracy and geometry, creativity and imagination, fine motor, gross motor, emotion recognition, empathy, communication, and collaboration. Scoring feeds radar charts, Shapiro-Wilk normality tests, ANOVA or Kruskal-Wallis tests, and post-hoc comparisons across settings.

What would settle it

Compare the LLM's per-setting ability scores for the same children against an independent, validated developmental assessment completed by trained observers who do not see the narratives; if the two rankings disagree or the setting differences disappear, the performance scores are measuring narrative content rather than development.

Watch

Extended reading notes

Core claim

The central claim is that LLM-based analysis of children's self-narratives is a reliable way to detect developmental abilities in free play. Using the qwen-max model with a structured prompt over eight ability categories, the approach achieved accuracy above 90 percent for identified abilities overall, with semantic consistency, ability relevance, and combined accuracy all high for cognitive, motor, and social domains. Emotional abilities scored lower, at roughly 70 to 90 percent, and the overall identification omission rate was 14.1 percent. The same pipeline produced performance scores that differed significantly across the four play settings for seven of eight abilities, with empathy showing no setting differences. The paper concludes that the approach is highly effective for identifying children's development across various free play settings.

Load-bearing premise

The whole result rests on treating the share of a child's play narratives in which the LLM infers an ability as a measure of how much the child has that ability; if children who talk more or whose activities are easier to label get higher scores without being more skilled, the setting differences describe the labeling process rather than development.

Editorial extensions

If this is right

  • Teachers can receive automatic, child-centred ability profiles from daily narratives, reducing reliance on memory and direct observation.
  • The four play areas show distinct developmental profiles, so a teacher can choose settings to target specific abilities, such as numeracy and geometry in the block area or gross motor skills on the hillside and playground.
  • Emotional abilities, especially empathy, are the least reliable outputs, so usable deployment should keep human verification for those domains.
  • Performance scores create longitudinal data that can support personalized learning plans and track changes over a semester.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Extending beyond the reported results, the performance score is a frequency of LLM inferences, so setting differences may partly capture what children choose to narrate or how easily activities in each area map to the ability labels, rather than true ability differences; comparing scores with an independent ability measure would separate these.
  • The same pipeline could be applied to other narrative sources, such as teacher observations or parent reports, to cross-validate the child's self-report.
  • Refining overlapping ability definitions, for example separating emotion recognition from empathy, is a plausible way to raise the emotional-domain accuracy the paper reports.
  • The absence of setting differences for empathy may mean empathy develops through stable relationships rather than in any single play area; a longer study or older age group could test this.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript proposes an LLM-based pipeline (qwen-max) that takes kindergarten children's self-narratives of free play, extracts behavior descriptions, and labels them with eight developmental abilities. A performance score for each child-ability-setting is defined as the number of narratives in which the LLM infers that ability divided by the number of activity records. Using 2,224 narratives from 29 children across four play areas, the authors report professional-rater agreement exceeding 90% on semantic consistency and ability relevance, and use Kruskal-Wallis/ANOVA to claim that play settings differ significantly in most ability dimensions. They conclude that the approach is 'highly effective' for identifying children's development across free play settings.

Significance. If valid, the approach would be practically useful: automated analysis of children's self-narratives could provide scalable, child-centered developmental feedback in early childhood education. The paper contributes a clearly described pipeline, a substantial longitudinal corpus of 2,224 naturally occurring play narratives, transparent prompt designs, and a professional-rater evaluation protocol. However, the significance depends entirely on whether the performance score measures child development. The current evidence supports only that expert raters find the LLM's narrative annotations plausible, not that the scores reflect actual abilities. The setting-comparison results are best interpreted as differences in LLM labeling rates rather than differences in developmental outcomes, which substantially limits the scientific contribution.

major comments (4)
  1. [Section 3.1; Section 5.1; Table 4] The central outcome variable is never validated as a developmental measure. In Section 3.1, Performance Scoring defines an ability score as the number of times the LLM infers that ability divided by the number of activity records, but no evidence is provided that this frequency is proportional to the child's actual level of the ability. The evaluation in Section 5.1 asks professionals only about Semantic Consistency and Ability Relevance, i.e., whether the LLM's behavior description matches the narrative text and whether the ability label matches that description. This establishes textual coherence, not criterion validity. Without comparison to an external criterion such as a standardized developmental assessment, teacher ratings, or direct observation, the high accuracy in Table 4 cannot support the claim that the scores measure development. The RQ2 comparisons therefore describe how often the LLM labels narratives from each setting, not how children develop there.
  2. [Section 4.2; Section 5.1; Table 5] The reliability evaluation has no inter-rater reliability and no independent ground truth. Section 4.2 states that each professional evaluated a random selection of samples 'without duplication,' so every item is rated by a single rater; no agreement statistic such as Cohen's or Fleiss' kappa can be computed, and the professional ratings are treated as error-free. In addition, the reported Accuracy in Table 4 is conditional on the LLM having identified an ability; the identification omissions in Table 5 (14.1% overall, and 26.5% for Gross Motor Development) are not incorporated into that accuracy. A fair overall performance measure must account for false negatives. The statement in Section 5.1.1 that the results represent 'high reliability' is therefore unsupported.
  3. [Section 5.2; Section 5.3; Table 7] The cross-setting comparisons in RQ2 are circular with respect to the LLM output. Table 7 reports significant Kruskal-Wallis/ANOVA differences across four play areas for seven of eight abilities, and Section 5.2 interprets these as showing that 'each play area may uniquely contribute' to children's development. But the dependent variable is the LLM inference-frequency score from Section 3.1. Differences in this score could arise from setting-specific narrative content (e.g., a zipline narrative naturally mentions climbing or running), LLM labeling biases, or differences in how children narrate in each area. The paper provides no control for narrative content and no external validation that these differences correspond to developmental outcomes. The conclusion in Section 6.1.2 that the approach 'can effectively reflect children's development' is thus not established.
  4. [Section 5.1.3; Section 6.3; Section 7] The paper's own evidence undercuts the strength of the conclusion. In Section 5.1.3, professional raters reported 'Misinterpretation of Activities and Abilities' and 'Overinterpretation and Subjectivity' as drawbacks, and Section 6.3 concedes that self-narratives alone may not capture the full range of developmental aspects. Table 5 shows omission rates above 20% for Numerical and Geometric Cognition, Gross Motor Development, and Communication. Section 7 nonetheless asserts that the approach is 'highly effective' and 'reliably' identifies children's performance. The conclusions should be tempered to match the evidence, or the paper should present additional validation data.
minor comments (5)
  1. [Section 4.3] The descriptive statistics 'mean is 76.67, variance is 13.05' are internally inconsistent with the reported minimum of 49 and maximum of 94; for 29 children, a variance of 13.05 is impossible given that range. The authors should report the correct dispersion statistic.
  2. [Section 5.2] The equal-interval segmentation thresholds (0.0-0.33 low, 0.34-0.66 moderate, 0.67-1.0 high) are introduced without justification; since the performance score is a frequency, the labels 'low/moderate/high' should be treated as descriptive conventions or derived from external benchmarks.
  3. [Section 4.3] The sample size determination is said to follow 'the standard formula,' but the parameters used (population proportion, margin of error, confidence level) are not reported; these should be stated for reproducibility.
  4. [Section 3.3; Table 1] The ability names in the prompt do not exactly match the names in Table 1 (e.g., 'Numerical and Geometric Cognition' vs 'Numeracy and Geometry,' and the prompt's list has inconsistent punctuation), which may contribute to labeling ambiguity and should be harmonized.
  5. [Table 4] The column header 'Count Total Consistency Semantic Consistency' appears to contain a typographical duplication; the intended header should be clarified.

Circularity Check

1 steps flagged · score 6.0 of 10

The paper defines 'performance score' as the rate at which the LLM infers an ability from a child's narratives, then reports cross-setting differences in that rate as 'developmental outcomes'; the central developmental claim is thus equivalent to the LLM's labeling behavior by construction.

  1. self definitional [Section 3.1 (Performance Scoring) and Abstract/Section 6.1.2 (RQ2)]
    "The calculation formula for a particular ability is the number of times that ability is inferred within a certain period, divided by the total number of activity records for that child during the same period. ... Moreover, significant differences in developmental outcomes were observed across play settings, highlighting each area’s unique contributions to specific abilities."

    The study's only developmental variable is the performance score defined in Section 3.1 as the frequency with which the LLM infers an ability from narratives. The abstract and Section 6.1.2 interpret differences in this variable as 'developmental outcomes' and 'distinct impacts on children's development.' Since the score is by construction the number of LLM inferences divided by activity records, the cross-setting comparisons are differences in LLM labeling rates. The professional-rater evaluation checks only semantic consistency and ability relevance of the LLM's outputs, not whether the score corresponds to any external measure of child ability; hence the claim that the approach reflects children's development reduces to the definition of the score.

full rationale

The paper is not circular at the level of text classification: Section 5.1 provides an independent human check that the LLM's ability labels and behavior descriptions are consistent with the narratives, and no parameters are fit to a target outcome. The circularity is in the developmental interpretation. The score used for RQ2 and the conclusion is defined, not measured against a criterion, as the LLM's inference frequency. 'Significant differences in developmental outcomes across play settings' are therefore significant differences in how often the LLM labels abilities in narratives from those settings. The paper's own limitation statement concedes that relying solely on self-narratives may be one-sided, but the deeper issue is that the score itself is not anchored to an independent assessment of child development. Because the central effectiveness claim is built on this self-defined score, the derivation is partially circular; however, the semantic-consistency evaluation gives the paper independent content, so the score is 6 rather than higher.

Assumptions & free parameters 1 free parameters · 4 assumptions · 2 invented entities

The central claim rests on an unvalidated scoring construct, a convenience gold standard of raters with no demonstrated reliability, and repeated-measures data treated as independent. The only fitted numeric parameter is the arbitrary segmentation threshold used for descriptive categories.

free parameters (1)
  • Equal-interval segmentation thresholds = 0.33 and 0.67
    Used in Section 5.2 to categorize scores as low, moderate, or high; chosen by hand, not derived from data.
assumptions (4)
  • domain assumption Children's self-narratives about play are valid indicators of their developmental abilities.
    The entire analysis uses narratives as the only data source; the paper itself notes in Section 6.3 that narratives may not capture the full range of development.
  • domain assumption The eight professionals' yes/no judgments are an accurate gold standard for ability identification.
    No inter-rater reliability, no training assessment, and no comparison with standardized developmental measures; the raters are introduced in Section 4.2.
  • ad hoc to paper Ability frequency, defined as LLM inference count divided by activity record count, is proportional to developmental performance.
    Defined in Section 3.1 with no empirical or theoretical justification that frequency equals ability level.
  • domain assumption Scores for the same child in different play areas can be treated as independent observations.
    Section 5.3 uses Kruskal-Wallis and ANOVA across areas while each child contributes a score to all four areas; repeated measures are ignored.
invented entities (2)
  • Eight-ability developmental framework (Table 1)
    purpose: Defines the target constructs the LLM must identify: numeracy, creativity, fine motor, gross motor, emotion recognition, empathy, communication, and collaboration.
    Adapted from ECDI2030 but redefined by the authors; no external validation that these categories measure child development.
  • Performance score based on LLM inference frequency
    purpose: Quantifies each child's ability level in each play area as the ratio of LLM inferences to activity records.
    The measure is constructed entirely from LLM outputs and has no independent calibration against standardized assessments.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Validating the Effectiveness of a Large Language Model-based Approach for Identifying Children's Development across Various Free Play Settings in Kindergarten." pith.science (2026). https://pith.science/paper/DIBBI3PM

@misc{pith2026250503369,
  author       = {Pith},
  title        = {Pith review of: Validating the Effectiveness of a Large Language Model-based Approach for Identifying Children's Development across Various Free Play Settings in Kindergarten},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DIBBI3PM}},
  note         = {Machine review of arXiv:2505.03369}
}
read the original abstract

Free play is a fundamental aspect of early childhood education, supporting children's cognitive, social, emotional, and motor development. However, assessing children's development during free play poses significant challenges due to the unstructured and spontaneous nature of the activity. Traditional assessment methods often rely on direct observations by teachers, parents, or researchers, which may fail to capture comprehensive insights from free play and provide timely feedback to educators. This study proposes an innovative approach combining Large Language Models (LLMs) with learning analytics to analyze children's self-narratives of their play experiences. The LLM identifies developmental abilities, while performance scores across different play settings are calculated using learning analytics techniques. We collected 2,224 play narratives from 29 children in a kindergarten, covering four distinct play areas over one semester. According to the evaluation results from eight professionals, the LLM-based approach achieved high accuracy in identifying cognitive, motor, and social abilities, with accuracy exceeding 90% in most domains. Moreover, significant differences in developmental outcomes were observed across play settings, highlighting each area's unique contributions to specific abilities. These findings confirm that the proposed approach is effective in identifying children's development across various free play settings. This study demonstrates the potential of integrating LLMs and learning analytics to provide child-centered insights into developmental trajectories, offering educators valuable data to support personalized learning and enhance early childhood education practices.

Figures

Figures reproduced from arXiv: 2505.03369 by the authors.

Figure 1
Figure 1. Technical framework of the LLM-base approach [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Children’s performance across different ability dimensions [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Differences in children’s performance across various ability dimensions in different settings [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

53 extracted references · 34 canonical work pages

  1. [1]

    Caillois, Man, play, and games, University of Illinois press, 2001

    R. Caillois, Man, play, and games, University of Illinois press, 2001

  2. [2]

    pdf, accessed January 9, 2025 (2012)

    Ministry of Education of China, Early learning and development guidelines for children aged 3 to 6 years, Retrieved from https: //www.unicef.cn/sites/unicef.org.china/files/2018-10/ 2012-national-early-learning-development-guidelines. pdf, accessed January 9, 2025 (2012)

  3. [3]

    Commonwealth of Australia, The early years learning framework for australia, Retrieved from https://www.acecqa.gov.au/sites/ default/files/2023-01/EYLF-2022-V2.0.pdf , accessed January 9, 2025 (2022)

  4. [4]

    Department for Education of UK, Early years foun- dation stage: Statutory framework, Retrieved from https://www.gov.uk/government/publications/ early-years-foundation-stage-framework--2 , accessed January 9, 2025 (2023)

  5. [5]

    Nordin, S

    N. Nordin, S. Mohamed, Exploring preschool teachers’ planning for the implementation of free play activities, International Journal of Aca- demic Research in Progressive Education and Development (2024).doi: https://doi.org/10.6007/ijarped/v13-i3/22239

  6. [6]

    Kurnia, S

    D. Kurnia, S. Winarni, S. Jarwo, G. F. Friskawati, ‘free play is impor- tant for children’s motor development, but how we can supervise it?’ a phenomenological study at early childhood education, Retos (2024). doi:https://doi.org/10.47197/retos.v58.104099

  7. [7]

    T. S. Haile, D. J. Ghirmai, Play-based learning: Conceptualization, benefits, and challenges of its implementation, European Scientific Journal, ESJ (2024). doi:https://doi.org/10.19044/esj.2024. v20n16p30

  8. [8]

    Achiam, S

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Ale- man, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al., Gpt-4 technical report, arXiv preprint arXiv:2303.08774 (2023). doi:https: //doi.org/10.48550/arXiv.2303.08774

Show all 53 references
  1. [9]

    Devlin, M.-W

    J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional transformers for language understanding, arXiv preprint arXiv:1810.04805 (2018). doi:https://doi.org/10. 48550/arXiv.1810.04805

  2. [10]

    J. Bai, S. Bai, Y . Chu, Z. Cui, K. Dang, X. Deng, Y . Fan, W. Ge, Y . Han, F. Huang, et al., Qwen technical report, arXiv preprint arXiv:2309.16609 (2023). doi:https://doi.org/10.48550/arXiv.2309.16609

  3. [11]

    Conneau, G

    A. Conneau, G. Lample, Cross-lingual language model pretraining, Ad- vances in neural information processing systems 32 (2019)

  4. [12]

    Rothe, S

    S. Rothe, S. Narayan, A. Severyn, Leveraging pre-trained checkpoints for sequence generation tasks, Transactions of the Association for Com- putational Linguistics 8 (2020) 264–280. doi:https://doi.org/10. 1162/tacl_a_00313

  5. [13]

    Zhang, S

    Y . Zhang, S. Sun, M. Galley, Y .-C. Chen, C. Brockett, X. Gao, J. Gao, J. Liu, B. Dolan, Dialogpt: Large-scale generative pre-training for con- versational response generation, arXiv preprint arXiv:1911.00536 (2019). doi:https://doi.org/10.48550/arXiv.1911.00536

  6. [14]

    Kasneci, K

    E. Kasneci, K. Seßler, S. K ¨uchemann, M. Bannert, D. Dementieva, F. Fis- cher, U. Gasser, G. Groh, S. G ¨unnemann, E. H ¨ullermeier, et al., Chat- gpt for good? on opportunities and challenges of large language models for education, Learning and individual differences 103 (20...

  7. [15]

    Abdelghani, Y .-H

    R. Abdelghani, Y .-H. Wang, X. Yuan, T. Wang, P. Lucas, H. Sauz´eon, P.- Y . Oudeyer, Gpt-3-driven pedagogical agents to train children’s curious question-asking skills, International Journal of Artificial Intelligence in Education 34 (2) (2024) 483–518. doi:https://doi.org/10...

  8. [16]

    Gabajiwala, P

    E. Gabajiwala, P. Mehta, R. Singh, R. Koshy, Quiz maker: Automatic quiz generation from text using nlp, in: Futuristic Trends in Networks and Computing Technologies: Select Proceedings of Fourth International Conference on FTNCT 2021, Springer, 2022, pp. 523–533. doi:https: //...

  9. [17]

    S. Kim, J. Shim, J. Shim, et al., A study on the utilization of openai chat- gpt as a second language learning tool, Journal of Multimedia Informa- tion System 10 (1) (2023) 79–88. doi:https://doi.org/10.33851/ JMIS.2023.10.1.79

  10. [18]

    M. Liu, L. J. Zhang, C. Biebricher, Investigating students’ cognitive pro- cesses in generative ai-assisted digital multimodal composing and tradi- tional writing, Computers & Education 211 (2024) 104977.doi:https: //doi.org/10.1016/j.compedu.2023.104977

  11. [19]

    Kohnke, B

    L. Kohnke, B. L. Moorhouse, D. Zou, Chatgpt for language teaching and learning, Relc Journal 54 (2) (2023) 537–550. doi:https://doi.org/ 10.1177/00336882231162868

  12. [20]

    Lan, N.-S

    Y .-J. Lan, N.-S. Chen, Teachers’ agency in the era of llm and generative ai, Educational Technology & Society 27 (1) (2024) I–XVIII

  13. [21]

    L. A. Razak, S. L. Yoong, J. Wiggers, P. J. Morgan, J. Jones, M. Finch, R. Sutherland, C. Lecathelnais, K. Gillham, T. Clinton- McHarg, et al., Impact of scheduling multiple outdoor free-play peri- ods in childcare on child moderate-to-vigorous physical activity: a clus- ter r...

  14. [23]

    Ellis, G

    C. Ellis, G. Beauchamp, S. Sarwar, J. Tyrie, D. Adams, S. Dumitrescu, C. Haughton, ‘oh no, the stick keeps falling!’: An analytical framework for conceptualising young children’s interactions during free play in a woodland setting, Journal of Early Childhood Research 19 (3) (2...

  15. [24]

    Tortella, M

    P. Tortella, M. Haga, J. Ingebrigtsen, G. Fumagalli, H. Sigmundsson, Comparing Free Play and Partly Structured Play in 4-5-Years-Old Chil- dren in an Outdoor Playground, FRONTIERS IN PUBLIC HEALTH 7 (2019). doi:https://doi.org/10.3389/fpubh.2019.00197

  16. [25]

    Tortella, M

    P. Tortella, M. Haga, H. Lors, G. F. Fumagalli, H. Sigmundsson, Effects of free play and partly structured playground activity on motor competence in preschool children: a pragmatic comparison trial, International Journal of Environmental Research and Public Health 19 (13) (20...

  17. [26]

    Palmer, K

    K. Palmer, K. Chinn, L. Robinson, The effect of the CHAMP intervention on fundamental motor skills and outdoor physical activity in preschoolers, Journal of Sport and Health Science 8 (2) (2019) 98–105. doi:https: //doi.org/10.1016/j.jshs.2018.12.003

  18. [27]

    Z. A. Jasem, D. Lambrick, D. C. Randall, A.-S. Darlington, The so- cial and physical environmental factors associated with the play of chil- dren living with life threatening/limiting conditions: Aq methodology study, Child: Care, Health and Development 48 (2) (2022) 336–346. ...

  19. [28]

    H. I. M. van Liempd, O. Oudgenoeg-Paz, R. G. Fukkink, P. P. Leseman, Young children’s exploration of the indoor playroom space in center- based childcare, Early Childhood Research Quarterly 43 (2018) 33–41. doi:https://doi.org/10.1016/j.ecresq.2017.11.005

  20. [29]

    S. Li, Q. Jiang, C. Deng, The Development and Validation of an Out- door Free Play Scale for Preschool Children, International Journal of Environmental Research and Public Health 20 (1) (2023). doi:https: //doi.org/10.3390/ijerph20010350

  21. [30]

    Tandon, K

    P. Tandon, K. Downing, B. Saelens, D. Christakis, Two Approaches to Increase Physical Activity for Preschool Children in Child Care Centers: A Matched-Pair Cluster-Randomized Trial, International Journal of En- vironmental Research and Public Health 16 (20) (2019). doi:https: ...

  22. [31]

    Ruiz-Esteban, J

    C. Ruiz-Esteban, J. Terry Andr ´es, I. M ´endez, ´A. Morales, Analysis of motor intervention program on the development of gross motor skills in preschoolers, International Journal of Environmental Research and Pub- lic Health 17 (13) (2020) 4891. doi:https://doi.org/10.3390/ ...

  23. [32]

    McCree, R

    M. McCree, R. Cutting, D. Sherwin, The hare and the tortoise go to forest school: taking the scenic route to academic attainment via emo- tional wellbeing outdoors, in: Young Children’s Emotional Experiences, Routledge, 2020, pp. 106–122. doi:https://doi.org/10.1080/ 09575146....

  24. [33]

    Kukkonen, S

    T. Kukkonen, S. Chang-Kredl, B. Bolden, Creative collaboration in young children’s playful group drawing, The Journal of creative behavior 54 (4) (2020) 897–911. doi:https://doi.org/10.1002/jocb.418

  25. [34]

    Verenikina, P

    I. Verenikina, P. Harris, P. Lysaght, Child’s play: computer games, theo- ries of play and children’s development, in: Proceedings of the interna- tional federation for information processing working group 3.5 open con- ference on Young children and learning technologies-V olu...

  26. [35]

    Colliver, L

    Y . Colliver, L. J. Harrison, J. E. Brown, P. Humburg, Free play pre- dicts self-regulation years later: Longitudinal evidence from a large aus- tralian sample of toddlers and preschoolers, Early Childhood Research Quarterly 59 (2022) 148–161. doi:https://doi.org/10.1016/j. ec...

  27. [36]

    A. Abe, R. Sanui, J. P. Loenneke, T. Abe, One-year handgrip strength change in kindergarteners depends upon physical activity sta- tus, Life 13 (8) (2023) 1665. doi:https://doi.org/10.3390/ life13081665

  28. [37]

    H. Q. Zeng, S. C. Ng, Free play matters: Promoting kindergarten chil- dren’s science learning using questioning strategies during loose parts play, Early Childhood Education Journal (2024) 1–16 doi:https:// doi.org/10.1007/s10643-024-01741-6

  29. [38]

    UNICEF, Early childhood development index 2030: A new tool to measure sdg indicator 4.2.1, Re- trieved from https://data.unicef.org/resources/ early-childhood-development-index-2030-ecdi2030/ , ac- cessed January 9, 2025 (2023)

  30. [39]

    B. Chen, Z. Zhang, N. Langren ´e, S. Zhu, Unleashing the potential of prompt engineering in large language models: a comprehensive review, arXiv preprint arXiv:2310.14735 (2023). doi:https://doi.org/10. 48550/arXiv.2310.14735

  31. [40]

    Sahoo, A

    P. Sahoo, A. K. Singh, S. Saha, V . Jain, S. Mondal, A. Chadha, A system- atic survey of prompt engineering in large language models: Techniques and applications, arXiv preprint arXiv:2402.07927 (2024). doi:https: //doi.org/10.48550/arXiv.2402.07927

  32. [41]

    openai.com/docs/guides/prompt-engineering/ six-strategies-for-getting-better-results/ , accessed January 9, 2025 (2023)

    OpenAI, Prompt engineering: Six strategies for get- ting better results, Retrieved from https://platform. openai.com/docs/guides/prompt-engineering/ six-strategies-for-getting-better-results/ , accessed January 9, 2025 (2023)

  33. [42]

    J. R. Coffino, C. Bailey, The anji play ecology of early learning, Child- hood Education 95 (1) (2019) 3–9. doi:https://doi.org/10.1080/ 00094056.2019.1565743

  34. [43]

    W. G. Cochran, Sampling techniques, john wiley & sons, 1977

  35. [44]

    Fleer, J

    M. Fleer, J. Cullen, A. Anning, Early Childhood education: Society and culture, SAGE, 2009

  36. [45]

    P. J. Hill, M. A. Mcnarry, K. A. Mackintosh, M. A. Murray, C. Pesce, N. C. Valentini, N. Getchell, P. D. Tomporowski, L. E. Robinson, L. M. Barnett, The influence of motor competence on broader aspects of health: a systematic review of the longitudinal associations between mot...

  37. [46]

    Sailunaz, M

    K. Sailunaz, M. Dhaliwal, J. Rokne, R. Alhajj, Emotion detection from text and speech: a survey, Social Network Analysis and Mining 8 (1) (2018) 28. doi:https://doi.org/10.1007/s13278-018-0505-2

  38. [47]

    Sachdeva, B

    N. Sachdeva, B. Coleman, W.-C. Kang, J. Ni, L. Hong, E. H. Chi, J. Caverlee, J. McAuley, D. Z. Cheng, How to train data-efficient llms, arXiv preprint arXiv:2402.09668 (2024). doi:https://doi.org/10. 48550/arXiv.2402.09668

  39. [48]

    Schroeder, Z

    K. Schroeder, Z. Wood-Doughty, Can you trust llm judgments? reliability of llm-as-a-judge, arXiv preprint arXiv:2412.12509 (2024).doi:https: //doi.org/10.48550/arXiv.2412.12509

  40. [49]

    Skulmowski, The cognitive architecture of digital externalization, Ed- ucational Psychology Review 35 (4) (2023) 101

    A. Skulmowski, The cognitive architecture of digital externalization, Ed- ucational Psychology Review 35 (4) (2023) 101. doi:https://doi. org/10.1007/s10648-023-09818-1

  41. [50]

    Garcia-Varela, Z

    F. Garcia-Varela, Z. Bekerman, M. Nussbaum, M. Mendoza, J. Mon- tero, Reducing interpretative ambiguity in an educational environment with chatgpt, Computers & Education 225 (2025) 105182. doi:https: //doi.org/10.1016/j.compedu.2024.105182

  42. [51]

    Y .-h. V . Chiang, M. Chang, N.-S. Chen, Can generative ai help realize the shift from an outcome-oriented to a process-outcome-balanced edu- cational practice?, Educational Technology & Society 27 (2) (2024) 347– 385

  43. [52]

    G. Siemens, Call for papers of the 1st international conference on learn- ing analytics & knowledge (lak 2011), in: Proceedings of the 1st Interna- tional Conference Learning Analytics & Knowledge, Banff, AL, Canada, V ol. 29, 2011, p. 2020

  44. [53]

    Alfirevi ´c, D

    N. Alfirevi ´c, D. Renduli ´c, M. Fo ˇsner, A. Fo ˇsner, Educational roles and scenarios for large language models: An ethnographic research study of artificial intelligence, in: Informatics, V ol. 11, MDPI, 2024, p. 78. doi: https://doi.org/10.3390/informatics11040078

  45. [54]

    Kaddour, J

    J. Kaddour, J. Harris, M. Mozes, H. Bradley, R. Raileanu, R. McHardy, Challenges and applications of large language models, arXiv preprint arXiv:2307.10169 (2023). doi:https://doi.org/10. 48550/arXiv.2307.10169

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.