REVIEW 3 major objections 5 minor 10 references
(AI peers) are people learning from the same standpoint: Perception of AI characters in a Collaborative Science Investigation
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This study claims that trust in AI characters is the prerequisite for social presence, and that social presence drives perceived effectiveness and the intention to adopt AI-led assessments.
desk verdict Useful early look at text-to-video AI characters in assessment, but the causal claim about trust and social presence overreaches the cross-sectional N=56 PLS-SEM design. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the PLS-SEM path model estimated on the 35 Likert-scale items, built on Kim et al.'s model of AI instructor trust and social presence. The model was 'optimized' from an initial exploratory model that had trust and social presence as simultaneous contributors to effectiveness; the optimized version places trust before social presence. The scenario itself, an SBA called SBLA-SP with a mentor (Ms. Rosie) and two teammates (Anika and Jimmy) delivered by text-to-video generated clips, provides the authentic collaborative context in which the perceptions are elicited. The four-point Likert subscales (social presence, trust, effectiveness, intention to adopt) and two open-ended questions are the measurement machinery.
What would settle it
A direct test would re-run the assessment with a larger sample and compare the optimized ordering against the reversed ordering (social presence to trust to effectiveness); if the reversed or a non-nested model fits the data as well or better, the claim that trust is a prerequisite for social presence fails. A manipulation that lowers trust, for example introducing a scripted factual error into the mentor's feedback in one condition, and observing whether perceived social presence drops accordingly would give causal evidence for the trust-to-presence leg.
Extended reading notes
Core claim
The paper's central claim is that learners' trust in AI characters shapes their sense of social presence, and social presence is what makes the characters feel effective; effectiveness then drives intention to adopt AI-driven learning. In the optimized PLS path model, trust predicts social presence at $\beta = 0.65$, social presence predicts perceived effectiveness at $\beta = 0.88$, and effectiveness predicts intention to adopt at $\beta = 0.78$, with all paths statistically significant. The authors interpret this as evidence that trust is a prerequisite for social presence in multimodal AI characters, and that no matter how lifelike the characters look, social presence cannot be fully realized unless trust is established. Qualitative responses add that trust is built from credible materials, reliable explanations and feedback, and consistent accuracy in information delivery.
Load-bearing premise
The causal ordering trust, social presence, effectiveness, intention to adopt is assumed rather than tested, and the final model was produced by optimizing the initial model after seeing the same 56 learners' cross-sectional self-reports.
Editorial extensions
If this is right
- Designers of multimodal AI characters for education should prioritize content credibility and goal alignment over visual realism, since trust is modeled as the entry point.
- Scenario-based assessments can use AI characters as both mentors and teammates without sacrificing perceived effectiveness, if the social-interactional context is preserved.
- Adoption of AI-driven learning tools may be predicted by perceived effectiveness, which itself depends on social presence, so usability metrics should track presence rather than only satisfaction.
- The finding extends prior voice-only AI instructor results to a multimodal, video-based setting, implying that synchronized visual and auditory output amplifies both the promise and the fragility of trust.
- Future studies using dynamically generated rather than pre-scripted characters will need safeguards such as retrieval-augmented generation and human oversight to maintain the trust on which the chain depends.
Reading between the lines
- If the ordering holds, interventions that increase trust (e.g., transparent error handling, verifiable sources) should have a larger effect on adoption than interventions that increase anthropomorphic fidelity.
- The 'peer-like' character framing may be doing separate work: learners' comments suggest that teammates who act from the same standpoint create attachment, hinting that role design interacts with social presence independently of content accuracy.
- The cross-sectional optimized model likely overstates path stability, so a preregistered replication with a confirmatory approach or a longitudinal design would be needed before the chain is treated as a causal law.
- For adaptive AI tutors, the model predicts that a single untrustworthy response could cascade through social presence and effectiveness, which is a testable design rule: reliability thresholds should be set before visual realism is increased.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a mixed-method study of 56 Korean high school English-as-a-second-language learners' perceptions of multimodal AI characters (a mentor and two teammates) embedded in a scenario-based assessment for collaborative science investigation. Quantitative data from 35 Likert-scale items covering technology self-efficacy, social presence, trust, effectiveness, and intention to adopt were analyzed with PLS-SEM, and two open-ended questions were thematically analyzed. The central quantitative claim is that learners' trust in the AI characters shaped their sense of social presence, which enhanced perceived effectiveness, which in turn drove intention to adopt AI-based learning tools; this path model is presented in Figure 4. The qualitative analysis identifies factors such as material credibility and alignment with learning goals as trust antecedents, and social presence as a key to perceived collaboration. The paper frames the study as an early exploration of AI characters in scenario-based assessment.
Significance. If the causal ordering were credible, the paper would make a useful contribution to the AIED literature by extending prior work on voice-only AI instructors to multimodal, video-based characters in scenario-based assessment, with practical implications for designing trustworthy AI characters. The study's strengths include a theory-grounded instrument with high overall internal consistency (Cronbach's alpha = .96), a concrete and authentic SBA context, clear construct definitions, and qualitative quotes that enrich the quantitative results. The use of PLS-SEM with bootstrapping is a standard choice for exploratory work, and the paper appropriately avoided reporting model fit indices that are unsuitable for PLS-SEM. However, the central causal claim is not supported by the design: the data are cross-sectional self-reports from a single assessment session, and the model was optimized post hoc on the same small sample, so the reported path coefficients are at risk of capitalizing on chance and cannot establish temporal precedence.
major comments (3)
- [Abstract, §4.2, §5.1] The central claim that 'learners' trust shaped their sense of social presence' is a causal claim, but the design cannot identify the direction of causation. The data are cross-sectional self-reports collected during a single assessment session (§3.3), and the path model is fit to these same data. PLS-SEM is a correlational technique and cannot distinguish trust→social presence from social presence→trust or from a common-cause explanation such as an overall positive attitude toward the AI characters. The paper itself describes the design as 'associational' (§3), yet §5.1 uses stronger language (trust as a 'prerequisite' and 'foundation'). The abstract and discussion should be revised to state that the model shows associations consistent with trust being related to social presence, and the causal ordering should be explicitly identified as an assumption or a direction for future longitudinal or experimental work.
- [§3.4, §4.2] The model optimization procedure is not reported transparently. The text states that 'The initial model was optimized' (§3.4) and then presents only the optimized model (§4.2, Figure 4), but it does not specify what modifications were made, how many alternative models were estimated, or what criteria drove the respecification. Because the final model and its path coefficients (e.g., β = 0.65 trust→social presence) are the product of selection on the same N=56 data, the reported coefficients are at risk of capitalizing on chance. The paper should provide a description of the specification search—for example, a table listing the initial model, the respecification steps, and the fit/reliability metrics used at each step—and should present the results as exploratory, noting that no holdout validation was performed.
- [§5.3] The limitations section omits the most serious inferential limitation: with cross-sectional, single-session self-report data and post hoc model optimization, the study cannot establish that trust precedes social presence or that these variables are causally related at all. Additionally, the sample size (N=56) with five latent constructs and 35 items is small for the complexity of the model; the '10-times rule' cited in §3.3 is a rule of thumb and does not address the stability of the specific path estimates or the risk of overfitting from model respecification. The limitation paragraph should acknowledge these points explicitly and recommend longitudinal or experimental validation.
minor comments (5)
- [§3.3] The survey was 'administered at three points during the test,' but the analysis appears to treat the responses as a single cross-section; please clarify whether the three administrations were pooled or analyzed separately, and if pooled, how within-subject dependence was handled.
- [§4.2] The phrase 'The moderate (social presence, r²=0.44) to high r² values (effectiveness, r²=0.76; intention to Adopt, r²=0.65)' mixes a vague qualitative label with precise values; please report the thresholds used for 'moderate' and 'high' and be consistent in capitalization ('Intention to Adopt').
- [Table 1] The table heading says k=37, and the rows list 35 Likert items plus 2 open-ended items; this is internally consistent, but the row 'Total Likert Scale Items 35' is redundant with the column totals and could be simplified to avoid confusion.
- [§2.2] The sentence 'Building on the advancement in multimodal AI, recent studies have e text-to-video...' appears to be missing a verb and should read 'have explored text-to-video' or similar.
- [Figure 4 caption] The caption states that 'dotted arrowed lines represent invalid paths,' but it is unclear whether these are non-significant paths, paths removed during optimization, or paths that were not included in the final model; please define the term 'invalid' explicitly.
Circularity Check
No significant circularity: the reported PLS-SEM paths are in-sample model estimates, not predictions derived from a fitted parameter renamed as a finding.
full rationale
The paper's central quantitative claim—that trust shaped social presence, which enhanced perceived effectiveness and intention to adopt—is an interpretive reading of path coefficients estimated from the same 56 self-report responses used for model optimization. That is a serious limitation for causal inference, but it is not circular under the stated criteria. Trust, social presence, and effectiveness are measured by separate Likert scales adapted from external sources (Kim et al. [7]; Choi & Ji [24]), so the constructs are not defined in terms of one another. The model optimization after data collection is explicitly disclosed in Section 3.4, and the reported coefficients are fitted associations rather than held-out predictions; the paper does not fit a parameter and then rename a close derivative as a predicted result. The self-citations ([13], [14]) support the broader SBLA-SP validation context and are not load-bearing for the specific trust→social presence→effectiveness path structure. The limitations section also candidly acknowledges the pre-scripted nature of the AI characters and the absence of learning-outcome measures. Consequently, although the causal ordering is under-identified and the small-sample post-hoc model is at risk of overfitting, no step in the derivation reduces by construction to its own inputs.
Assumptions & free parameters
free parameters (4)
- Path coefficient: trust to social presence =
beta = 0.65
- Path coefficient: social presence to effectiveness =
beta = 0.88
- Path coefficient: effectiveness to intention to adopt =
beta = 0.78
- Factor loadings for 35 Likert items =
not reported in preprint
assumptions (4)
- domain assumption Survey items adapted from Kim et al. [7], Choi and Ji [24], and Nass et al. [25] validly measure the named constructs.
- domain assumption PLS-SEM is appropriate for N=56 and the 10-times rule is sufficient.
- domain assumption Common method bias does not distort the path estimates.
- domain assumption Deductive coding of open-ended responses is reliable.
Cite this review
Pith. "Pith review of (AI peers) are people learning from the same standpoint: Perception of AI characters in a Collaborative Science Investigation." pith.science (2026). https://pith.science/paper/MZI7YY2X
@misc{pith2026250606165,
author = {Pith},
title = {Pith review of: (AI peers) are people learning from the same standpoint: Perception of AI characters in a Collaborative Science Investigation},
year = {2026},
howpublished = {\url{https://pith.science/paper/MZI7YY2X}},
note = {Machine review of arXiv:2506.06165}
}
read the original abstract
While the complexity of 21st-century demands has promoted pedagogical approaches to foster complex competencies, a persistent gap remains between in-class learning activities and individualized learning or assessment practices. To address this, studies have explored the use of AI-generated characters in learning and assessment. One attempt is scenario-based assessment (SBA), a technique that not only measures but also fosters the development of competencies throughout the assessment process. SBA introduces simulated agents to provide an authentic social-interactional context, allowing for the assessment of competency-based constructs while mitigating the unpredictability of real-life interactions. Recent advancements in multimodal AI, such as text-to-video technology, allow these agents to be enhanced into AI-generated characters. This mixed-method study investigates how learners perceive AI characters taking the role of mentor and teammates in an SBA mirroring the context of a collaborative science investigation. Specifically, we examined the Likert scale responses of 56 high schoolers regarding trust, social presence, and effectiveness. We analyzed the relationships between these factors and their impact on the intention to adopt AI characters through PLS-SEM. Our findings indicated that learners' trust shaped their sense of social presence with the AI characters, enhancing perceived effectiveness. Qualitative analysis further highlighted factors that foster trust, such as material credibility and alignment with learning goals, as well as the pivotal role of social presence in creating a collaborative context. This paper was accepted as an full paper for AIED 2025.
Reference graph
Works this paper leans on
-
[1]
OECD. (2015). PISA 2015 Collaborative problem-solving framework . 2. Mislevy, R. J. (2018). Sociocognitive foundations of educational measurement . Routledge
work page 2015
-
[3]
von Davier, A. A. (2017). Computational psychometrics in support of collaborative educational assessments. Journal of Educational Measurement, 54 (1), 3-11
work page 2017
-
[4]
Pataranutaporn, P., Danry, V., Leong, J., Punpongsanon, P., Novy, D., Maes, P., & Sra, M. (2021). AI-generated characters for supporting personalized learning and well-being. Nature Machine Intelligence, 3 (12), 1013–1022. https://doi.org/10.1038/s42256-021-00417-9 5. Edwards, C., Edwards, A., Stoll, B., Lin, X., & Massey, N. (2019). Evaluations of an art...
arXiv 2021
-
[10]
Banerjee, H. L. (2019). Investigating the construct of topical knowledge in second language assessment: A scenario-based assessment approach. Language Assessment Quarterly, 16 (2), 133-160
work page 2019
-
[11]
Bennett, R.E. (2015). CBAL: Results from piloting innovative K-12 assessments (Research Report No. RR-11-23). Educational Testing Service
work page 2015
-
[12]
Purpura, J. E. (2021). A rationale for using a scenario-based assessment to measure competency-based, situated second and foreign language proficiency. In M. Masperi, C. Cervini, & Y. Bardière (Eds.), Évaluation des acquisitions langagières: Du formatif au certificatif . MediAzioni, 32 , A54-A96. http://www.mediazioni.sitlec.unibo.it 13. Purpura, J. E., J...
work page 2021
-
[14]
Voss, E., Joo, S-H., Eskin, D., & Vafaee, P. (2025). A Learning-Oriented Language Assessment (LOLA) framework argument for artificial intelligence (AI) technologies in scenario-based language assessments (SBLA) Language Assessment Quarterly
work page 2025
-
[15]
Qin, F., Li, K., & Yan, J. (2020). Understanding user trust in artificial intelligence-based educational systems: Evidence from China. British Journal of Educational Technology, 51 (5), 1693–1710. https://doi.org/10.1111/bjet.12994 16. Amoozadeh, M., Daniels, D., Nam, D., Kumar, A., Chen, S., Hilton, M., Srinivasa Ragavan, S., & Alipour, M. A. (2024). Tru...
arXiv 2020
Show all 10 references
-
[21]
F., Risher, J
Hair, J. F., Risher, J. J., Sarstedt, M., & Ringle, C. M. (2019). When to use and how to report the results of PLS-SEM. European Business Review, 31 (1), 2-24
2019
-
[22]
R., Rieser, V., & Gabriel, I
Manzini, A., Keeling, G., Marchal, N., McKee, K. R., Rieser, V., & Gabriel, I. (2024). Should users trust advanced AI assistants? Justified trust as a function of competence and alignment. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , 1...
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.