Pith. sign in

REVIEW 3 major objections 5 minor 10 references

(AI peers) are people learning from the same standpoint: Perception of AI characters in a Collaborative Science Investigation

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This study claims that trust in AI characters is the prerequisite for social presence, and that social presence drives perceived effectiveness and the intention to adopt AI-led assessments.

desk verdict Useful early look at text-to-video AI characters in assessment, but the causal claim about trust and social presence overreaches the cross-sectional N=56 PLS-SEM design. read the letter →

arxiv 2506.06165 v1 pith:MZI7YY2X submitted 2025-06-06 cs.HC cs.AI

classification cs.HCcs.AI
keywords AI-generatedcharactersscenario-basedassessmentsocialpresencetrustPLS-SEMcollaborativescienceinvestigationtext-to-videointentiontoadopt
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper investigates how high school learners perceive AI-generated characters that play the roles of mentor and teammates in a scenario-based science investigation. Using survey responses from 56 Korean learners learning English as a second language, it argues that trust in the AI characters is the foundation that enables a sense of social presence, that social presence drives perceived effectiveness, and that perceived effectiveness leads to intention to adopt AI characters in future learning. The authors position this trust-to-presence-to-effectiveness chain as the mechanism by which multimodal AI characters can support rather than distract from collaborative assessment. The study matters because scenario-based assessments are being enhanced with text-to-video AI characters, and it claims to identify which perception to build first.

What carries the argument

The load-bearing object is the PLS-SEM path model estimated on the 35 Likert-scale items, built on Kim et al.'s model of AI instructor trust and social presence. The model was 'optimized' from an initial exploratory model that had trust and social presence as simultaneous contributors to effectiveness; the optimized version places trust before social presence. The scenario itself, an SBA called SBLA-SP with a mentor (Ms. Rosie) and two teammates (Anika and Jimmy) delivered by text-to-video generated clips, provides the authentic collaborative context in which the perceptions are elicited. The four-point Likert subscales (social presence, trust, effectiveness, intention to adopt) and two open-ended questions are the measurement machinery.

What would settle it

A direct test would re-run the assessment with a larger sample and compare the optimized ordering against the reversed ordering (social presence to trust to effectiveness); if the reversed or a non-nested model fits the data as well or better, the claim that trust is a prerequisite for social presence fails. A manipulation that lowers trust, for example introducing a scripted factual error into the mentor's feedback in one condition, and observing whether perceived social presence drops accordingly would give causal evidence for the trust-to-presence leg.

Watch

Extended reading notes

Core claim

The paper's central claim is that learners' trust in AI characters shapes their sense of social presence, and social presence is what makes the characters feel effective; effectiveness then drives intention to adopt AI-driven learning. In the optimized PLS path model, trust predicts social presence at $\beta = 0.65$, social presence predicts perceived effectiveness at $\beta = 0.88$, and effectiveness predicts intention to adopt at $\beta = 0.78$, with all paths statistically significant. The authors interpret this as evidence that trust is a prerequisite for social presence in multimodal AI characters, and that no matter how lifelike the characters look, social presence cannot be fully realized unless trust is established. Qualitative responses add that trust is built from credible materials, reliable explanations and feedback, and consistent accuracy in information delivery.

Load-bearing premise

The causal ordering trust, social presence, effectiveness, intention to adopt is assumed rather than tested, and the final model was produced by optimizing the initial model after seeing the same 56 learners' cross-sectional self-reports.

Editorial extensions

If this is right

  • Designers of multimodal AI characters for education should prioritize content credibility and goal alignment over visual realism, since trust is modeled as the entry point.
  • Scenario-based assessments can use AI characters as both mentors and teammates without sacrificing perceived effectiveness, if the social-interactional context is preserved.
  • Adoption of AI-driven learning tools may be predicted by perceived effectiveness, which itself depends on social presence, so usability metrics should track presence rather than only satisfaction.
  • The finding extends prior voice-only AI instructor results to a multimodal, video-based setting, implying that synchronized visual and auditory output amplifies both the promise and the fragility of trust.
  • Future studies using dynamically generated rather than pre-scripted characters will need safeguards such as retrieval-augmented generation and human oversight to maintain the trust on which the chain depends.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the ordering holds, interventions that increase trust (e.g., transparent error handling, verifiable sources) should have a larger effect on adoption than interventions that increase anthropomorphic fidelity.
  • The 'peer-like' character framing may be doing separate work: learners' comments suggest that teammates who act from the same standpoint create attachment, hinting that role design interacts with social presence independently of content accuracy.
  • The cross-sectional optimized model likely overstates path stability, so a preregistered replication with a confirmatory approach or a longitudinal design would be needed before the chain is treated as a causal law.
  • For adaptive AI tutors, the model predicts that a single untrustworthy response could cascade through social presence and effectiveness, which is a testable design rule: reliability thresholds should be set before visual realism is increased.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports a mixed-method study of 56 Korean high school English-as-a-second-language learners' perceptions of multimodal AI characters (a mentor and two teammates) embedded in a scenario-based assessment for collaborative science investigation. Quantitative data from 35 Likert-scale items covering technology self-efficacy, social presence, trust, effectiveness, and intention to adopt were analyzed with PLS-SEM, and two open-ended questions were thematically analyzed. The central quantitative claim is that learners' trust in the AI characters shaped their sense of social presence, which enhanced perceived effectiveness, which in turn drove intention to adopt AI-based learning tools; this path model is presented in Figure 4. The qualitative analysis identifies factors such as material credibility and alignment with learning goals as trust antecedents, and social presence as a key to perceived collaboration. The paper frames the study as an early exploration of AI characters in scenario-based assessment.

Significance. If the causal ordering were credible, the paper would make a useful contribution to the AIED literature by extending prior work on voice-only AI instructors to multimodal, video-based characters in scenario-based assessment, with practical implications for designing trustworthy AI characters. The study's strengths include a theory-grounded instrument with high overall internal consistency (Cronbach's alpha = .96), a concrete and authentic SBA context, clear construct definitions, and qualitative quotes that enrich the quantitative results. The use of PLS-SEM with bootstrapping is a standard choice for exploratory work, and the paper appropriately avoided reporting model fit indices that are unsuitable for PLS-SEM. However, the central causal claim is not supported by the design: the data are cross-sectional self-reports from a single assessment session, and the model was optimized post hoc on the same small sample, so the reported path coefficients are at risk of capitalizing on chance and cannot establish temporal precedence.

major comments (3)
  1. [Abstract, §4.2, §5.1] The central claim that 'learners' trust shaped their sense of social presence' is a causal claim, but the design cannot identify the direction of causation. The data are cross-sectional self-reports collected during a single assessment session (§3.3), and the path model is fit to these same data. PLS-SEM is a correlational technique and cannot distinguish trust→social presence from social presence→trust or from a common-cause explanation such as an overall positive attitude toward the AI characters. The paper itself describes the design as 'associational' (§3), yet §5.1 uses stronger language (trust as a 'prerequisite' and 'foundation'). The abstract and discussion should be revised to state that the model shows associations consistent with trust being related to social presence, and the causal ordering should be explicitly identified as an assumption or a direction for future longitudinal or experimental work.
  2. [§3.4, §4.2] The model optimization procedure is not reported transparently. The text states that 'The initial model was optimized' (§3.4) and then presents only the optimized model (§4.2, Figure 4), but it does not specify what modifications were made, how many alternative models were estimated, or what criteria drove the respecification. Because the final model and its path coefficients (e.g., β = 0.65 trust→social presence) are the product of selection on the same N=56 data, the reported coefficients are at risk of capitalizing on chance. The paper should provide a description of the specification search—for example, a table listing the initial model, the respecification steps, and the fit/reliability metrics used at each step—and should present the results as exploratory, noting that no holdout validation was performed.
  3. [§5.3] The limitations section omits the most serious inferential limitation: with cross-sectional, single-session self-report data and post hoc model optimization, the study cannot establish that trust precedes social presence or that these variables are causally related at all. Additionally, the sample size (N=56) with five latent constructs and 35 items is small for the complexity of the model; the '10-times rule' cited in §3.3 is a rule of thumb and does not address the stability of the specific path estimates or the risk of overfitting from model respecification. The limitation paragraph should acknowledge these points explicitly and recommend longitudinal or experimental validation.
minor comments (5)
  1. [§3.3] The survey was 'administered at three points during the test,' but the analysis appears to treat the responses as a single cross-section; please clarify whether the three administrations were pooled or analyzed separately, and if pooled, how within-subject dependence was handled.
  2. [§4.2] The phrase 'The moderate (social presence, r²=0.44) to high r² values (effectiveness, r²=0.76; intention to Adopt, r²=0.65)' mixes a vague qualitative label with precise values; please report the thresholds used for 'moderate' and 'high' and be consistent in capitalization ('Intention to Adopt').
  3. [Table 1] The table heading says k=37, and the rows list 35 Likert items plus 2 open-ended items; this is internally consistent, but the row 'Total Likert Scale Items 35' is redundant with the column totals and could be simplified to avoid confusion.
  4. [§2.2] The sentence 'Building on the advancement in multimodal AI, recent studies have e text-to-video...' appears to be missing a verb and should read 'have explored text-to-video' or similar.
  5. [Figure 4 caption] The caption states that 'dotted arrowed lines represent invalid paths,' but it is unclear whether these are non-significant paths, paths removed during optimization, or paths that were not included in the final model; please define the term 'invalid' explicitly.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the reported PLS-SEM paths are in-sample model estimates, not predictions derived from a fitted parameter renamed as a finding.

full rationale

The paper's central quantitative claim—that trust shaped social presence, which enhanced perceived effectiveness and intention to adopt—is an interpretive reading of path coefficients estimated from the same 56 self-report responses used for model optimization. That is a serious limitation for causal inference, but it is not circular under the stated criteria. Trust, social presence, and effectiveness are measured by separate Likert scales adapted from external sources (Kim et al. [7]; Choi & Ji [24]), so the constructs are not defined in terms of one another. The model optimization after data collection is explicitly disclosed in Section 3.4, and the reported coefficients are fitted associations rather than held-out predictions; the paper does not fit a parameter and then rename a close derivative as a predicted result. The self-citations ([13], [14]) support the broader SBLA-SP validation context and are not load-bearing for the specific trust→social presence→effectiveness path structure. The limitations section also candidly acknowledges the pre-scripted nature of the AI characters and the absence of learning-outcome measures. Consequently, although the causal ordering is under-identified and the small-sample post-hoc model is at risk of overfitting, no step in the derivation reduces by construction to its own inputs.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

No new theoretical entities are introduced; the AI characters are existing technology. The central quantitative claim depends on estimated path coefficients and factor loadings from a small single-sample PLS-SEM model. The qualitative coding and construct validity rest on domain assumptions that are only partially documented.

free parameters (4)
  • Path coefficient: trust to social presence = beta = 0.65
    Estimated from N=56 survey data in the optimized PLS-SEM model (Section 4.2, Figure 4).
  • Path coefficient: social presence to effectiveness = beta = 0.88
    Estimated from N=56 survey data in the optimized PLS-SEM model (Section 4.2, Figure 4).
  • Path coefficient: effectiveness to intention to adopt = beta = 0.78
    Estimated from N=56 survey data in the optimized PLS-SEM model (Section 4.2, Figure 4).
  • Factor loadings for 35 Likert items = not reported in preprint
    Measurement model in Appendix A is referenced but not included, so the loadings used to construct latent variables cannot be inspected.
assumptions (4)
  • domain assumption Survey items adapted from Kim et al. [7], Choi and Ji [24], and Nass et al. [25] validly measure the named constructs.
    The paper does not list the items or report validation statistics beyond Cronbach's alpha; Appendix A is referenced but absent.
  • domain assumption PLS-SEM is appropriate for N=56 and the 10-times rule is sufficient.
    Section 3.3 states sample size meets the 10-times rule; this is a contested rule of thumb for SEM and is not a replacement for full power analysis.
  • domain assumption Common method bias does not distort the path estimates.
    All constructs are self-report Likert items collected at the same time from the same participants; the paper says common method bias was evaluated in Appendix A, which is not shown.
  • domain assumption Deductive coding of open-ended responses is reliable.
    Section 4.1 describes deductive coding with social presence, trust, and effectiveness as pillars, but no inter-rater reliability or codebook is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of (AI peers) are people learning from the same standpoint: Perception of AI characters in a Collaborative Science Investigation." pith.science (2026). https://pith.science/paper/MZI7YY2X

@misc{pith2026250606165,
  author       = {Pith},
  title        = {Pith review of: (AI peers) are people learning from the same standpoint: Perception of AI characters in a Collaborative Science Investigation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MZI7YY2X}},
  note         = {Machine review of arXiv:2506.06165}
}
read the original abstract

While the complexity of 21st-century demands has promoted pedagogical approaches to foster complex competencies, a persistent gap remains between in-class learning activities and individualized learning or assessment practices. To address this, studies have explored the use of AI-generated characters in learning and assessment. One attempt is scenario-based assessment (SBA), a technique that not only measures but also fosters the development of competencies throughout the assessment process. SBA introduces simulated agents to provide an authentic social-interactional context, allowing for the assessment of competency-based constructs while mitigating the unpredictability of real-life interactions. Recent advancements in multimodal AI, such as text-to-video technology, allow these agents to be enhanced into AI-generated characters. This mixed-method study investigates how learners perceive AI characters taking the role of mentor and teammates in an SBA mirroring the context of a collaborative science investigation. Specifically, we examined the Likert scale responses of 56 high schoolers regarding trust, social presence, and effectiveness. We analyzed the relationships between these factors and their impact on the intention to adopt AI characters through PLS-SEM. Our findings indicated that learners' trust shaped their sense of social presence with the AI characters, enhancing perceived effectiveness. Qualitative analysis further highlighted factors that foster trust, such as material credibility and alignment with learning goals, as well as the pivotal role of social presence in creating a collaborative context. This paper was accepted as an full paper for AIED 2025.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

10 extracted references · 7 canonical work pages

  1. [1]

    OECD. (2015). PISA 2015 Collaborative problem-solving framework . 2. Mislevy, R. J. (2018). Sociocognitive foundations of educational measurement . Routledge

  2. [3]

    von Davier, A. A. (2017). Computational psychometrics in support of collaborative educational assessments. Journal of Educational Measurement, 54 (1), 3-11

  3. [4]

    Pataranutaporn, P., Danry, V., Leong, J., Punpongsanon, P., Novy, D., Maes, P., & Sra, M. (2021). AI-generated characters for supporting personalized learning and well-being. Nature Machine Intelligence, 3 (12), 1013–1022. https://doi.org/10.1038/s42256-021-00417-9 5. Edwards, C., Edwards, A., Stoll, B., Lin, X., & Massey, N. (2019). Evaluations of an art...

  4. [10]

    Banerjee, H. L. (2019). Investigating the construct of topical knowledge in second language assessment: A scenario-based assessment approach. Language Assessment Quarterly, 16 (2), 133-160

  5. [11]

    Bennett, R.E. (2015). CBAL: Results from piloting innovative K-12 assessments (Research Report No. RR-11-23). Educational Testing Service

  6. [12]

    Purpura, J. E. (2021). A rationale for using a scenario-based assessment to measure competency-based, situated second and foreign language proficiency. In M. Masperi, C. Cervini, & Y. Bardière (Eds.), Évaluation des acquisitions langagières: Du formatif au certificatif . MediAzioni, 32 , A54-A96. http://www.mediazioni.sitlec.unibo.it 13. Purpura, J. E., J...

  7. [14]

    Voss, E., Joo, S-H., Eskin, D., & Vafaee, P. (2025). A Learning-Oriented Language Assessment (LOLA) framework argument for artificial intelligence (AI) technologies in scenario-based language assessments (SBLA) Language Assessment Quarterly

  8. [15]

    Qin, F., Li, K., & Yan, J. (2020). Understanding user trust in artificial intelligence-based educational systems: Evidence from China. British Journal of Educational Technology, 51 (5), 1693–1710. https://doi.org/10.1111/bjet.12994 16. Amoozadeh, M., Daniels, D., Nam, D., Kumar, A., Chen, S., Hilton, M., Srinivasa Ragavan, S., & Alipour, M. A. (2024). Tru...

Show all 10 references
  1. [21]

    F., Risher, J

    Hair, J. F., Risher, J. J., Sarstedt, M., & Ringle, C. M. (2019). When to use and how to report the results of PLS-SEM. European Business Review, 31 (1), 2-24

  2. [22]

    R., Rieser, V., & Gabriel, I

    Manzini, A., Keeling, G., Marchal, N., McKee, K. R., Rieser, V., & Gabriel, I. (2024). Should users trust advanced AI assistants? Justified trust as a function of competence and alignment. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , 1...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.