REVIEW 3 major objections 6 minor 40 references
Even with repeated real exam scores as anchors, local peer talk under incomplete social views can steadily warp a class’s shared sense of who ranks high.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 21:16 UTC pith:2DA3ABV2
load-bearing objection Real-classroom multi-agent sim with a clean DPAE rise under local visibility; useful as mechanism work, but ablations undercut the causal story the abstract sells. the 3 major comments →
Subjective-Graph LLM Agents for Simulating Uncertainty in Classroom Social Perception
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Graph-local evidence and credibility-weighted communication among LLM agents, each confined to an individualized subjective social graph, can generate persistent and accumulating distortions in perceived academic standing even when agents repeatedly receive objective exam scores as self-anchors. On the Social-Observed subset of 419 students in twelve classrooms, group ranking error rises from 0.066 ± 0.008 to 0.124 ± 0.009 over six epochs, while the framework avoids the near-total opinion collapse seen in evaluated DeGroot configurations.
What carries the argument
Individualized subjective graphs: questionnaire-derived personal networks that gate retrieval-augmented evidence, interaction partners, and reachability, combined with LLM trust gating and precision-weighted Bayesian fusion of Gaussian peer-ability beliefs.
Load-bearing premise
The central premise is that language-model assessments and trust scores, run on questionnaire-built personal graphs with parametric noise for social anxiety, are good enough stand-ins for how real students judge peers—without checking simulated peer rankings against students’ own reported rankings of classmates.
What would settle it
Re-run the same six-epoch protocol and check whether group ranking error still rises as reported; or collect independent student-reported peer rankings in the same classrooms and test whether the simulated beliefs track those reports rather than only the exam-score ranking.
If this is right
- Collective misperception of academic rank can grow across successive exams even when every student keeps receiving an objective self-score.
- Heterogeneous local visibility and explicit trust gating matter more for long-horizon ranking fidelity than unconstrained global retrieval.
- Classical consensus dynamics that drive opinion diversity near zero can look competitive on correlation while erasing useful disagreement.
- Classrooms can differ systematically in how fast misperception accumulates, depending on subjective network structure.
- Language models can be used as structurally constrained cognitive modules rather than omniscient oracles in multi-agent social simulation.
Where Pith is reading between the lines
- If the mechanism holds, interventions that expand trusted local visibility may reduce ranking misperception more than simply giving more frequent grades.
- The same subjective-graph design is a natural stress test for adult workplaces or online cohorts where sparse social data and performance signals co-exist.
- Comparing simulated peer rankings to students’ actual questionnaire rankings of classmates would most cleanly separate cognitive realism from fitting exam ranks alone.
- Preserving opinion diversity may be as useful a diagnostic as ranking error when evaluating social simulators.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a multi-agent simulation framework in which LLM agents, each restricted to an individualized subjective social graph built from questionnaire data, exchange uncertainty-annotated peer assessments, apply LLM-derived trust weights, and maintain Gaussian beliefs updated by precision-weighted Bayesian fusion. Using 12 middle-school classrooms (482 students; Social-Observed n=419) and six consecutive exams as exogenous self-anchors, the authors report that collective ranking error (DPAE) rises from 0.066±0.008 to 0.124±0.009 despite repeated score anchoring. Ablations remove constrained RAG, shared vs individualized visibility, and LLM trust; external baselines include feature regressors, simple graph aggregators, and DeGroot dynamics. The stated contribution is mechanism-oriented: showing that graph-local evidence and credibility-weighted communication can sustain structured misperception without global visibility, while preserving more opinion diversity than consensus-collapsing DeGroot runs.
Significance. If the results hold under tighter causal and external checks, the work is a useful methodological contribution at the intersection of multi-agent systems, epistemic uncertainty, and educational social networks. Strengths include real multi-classroom longitudinal scores, explicit probabilistic belief states (Eqs. 2–3), multi-seed reporting with means±std, ablations, DeGroot diversity comparisons, and released code. The framing as a constrained simulation of perception—not a stronger rank predictor—is appropriate. The main scientific value is a concrete architecture for embedding LLMs under local visibility and Bayesian fusion rather than a new empirical law of classroom psychology. That value is reduced until mechanism attribution and external face validity of the simulated perceptions are clarified.
major comments (3)
- [§IV.C, Table I; Abstract] The central claim that graph-local evidence and credibility-weighted communication generate the reported DPAE trajectory is only weakly supported by Table I. No-RAG reaches DPAE6=0.121±0.011 and ρ6=0.879, matching or slightly beating Baseline (0.124±0.009, 0.876), while No-Subjective-Graph and No-LLM-Trust produce only modest degradations (0.128 and 0.131). The abstract and §IV.C still credit individualized visibility and LLM trust for “more stable long-horizon behavior” and treat constrained retrieval mainly as an anti-leakage safeguard. That reading is not forced by the numbers: the accumulation of ranking error appears largely driven by repeated noisy narrative messaging plus Gaussian fusion under some visibility, not specifically by subjective-graph-constrained RAG. A load-bearing control is missing—e.g., score-only or frozen-message communication, or a non-LLM observation channel—so
- [§III.A–C; §IV.B; Abstract/Title] All evaluation is against exam-induced ground-truth ranks R(t) (Eq. 6 and §IV.A–B). The paper never compares simulated peer rankings or belief means to students’ actual reported rankings or competence judgments of classmates, even though the questionnaire supplies social-relational and psychometric inputs used to build G_i and α_i (§III.A). The title, abstract, and introduction frame the object of study as classroom social perception and subjective academic standing. Under that framing, agreement with exam ranks alone shows that the simulation can drift from objective performance; it does not show that the drift resembles real student misperception. At minimum, the manuscript should either (i) report any available peer-perception items as external validation, or (ii) substantially narrow claims from “social perception” to “simulated ranking beliefs under local interaction,” and discuss t
- [§IV.D, Tables II–III; Abstract] The DeGroot comparison in §IV.D / Tables II–III is informative on diversity collapse but incomplete as a ranking baseline. One-step DeGroot achieves DPAE6=0.048 and Acc@3=0.528, and five-step DeGroot still has DPAE6=0.103 with Acc@3=0.444—both better ranking metrics than the proposed method’s Epoch-6 DPAE 0.124 and Acc@3 0.278—while the paper highlights the 30-step configuration (ρ6=0.846, Div6≈5×10−4) as the main foil. The diversity–accuracy trade-off is real and worth reporting, but the claim that the framework “achieves lower final ranking error” than “evaluated DeGroot configurations” (Abstract) depends on which step count is treated as the comparator. Please report the full step sweep as primary baselines, state the selection rule for the DeGroot configuration used in headline comparisons, and qualify “lower final ranking error” accordingly rather than implying a uniform win over De
minor comments (6)
- [Title; Abstract; Fig. 1] Title/abstract use “Subjective-Graph LLM Agents…” while the body and Fig. 1 use “Rashomon Set Agents.” Align naming throughout for indexing and clarity.
- [§III.C, Eq. (3)] Eq. (3) writes the Bayesian update without time/round superscripts on μ, σ, τ, ŝ; the surrounding text uses (t,r). Make the notation consistent.
- [§IV.A Protocol; §III.A] Free parameters R=2, K=3, σ_self, and the mapping from questionnaire scales to α_i are stated but not sensitivity-tested beyond the three ablations. A short appendix sweep (or explicit fixed defaults with justification) would help reproducibility.
- [§IV.A Metrics; Table I–II] Fig. 2–8 captions are clear, but Acc@3 is reported without a chance baseline or class-size normalization; with varying classroom sizes, raw Top-3 accuracy is hard to interpret across classes.
- [§II Related Work] Related Work cites DeGroot, cognitive social structures (Krackhardt), and LLM multi-agent surveys appropriately; a brief pointer to empirical classroom peer-perception / sociometric ranking studies would better situate the missing external validation.
- [Abstract; throughout] Minor prose issues: “W e” / “T o” spacing artifacts in the abstract; “difficult”/“insufficiently” ligature rendering; “Full-Temporal” hyphenation inconsistency.
Circularity Check
No circular derivation: rising DPAE is a simulation outcome under external exam anchors and baselines, not an identity or fitted-parameter rename.
full rationale
The paper is a mechanism-oriented multi-agent simulation, not a first-principles derivation of a closed-form prediction. Exam scores Y^(t) enter only as exogenous self-anchors (Eq. 4) and as the evaluation target R^(t) for DPAE = 1 − Spearman(ˆR, R); peer beliefs are Gaussian states updated by LLM messages and trust-weighted Bayesian fusion (Eqs. 2–3) under questionnaire-built subjective graphs, not by fitting parameters to the DPAE trajectory. Nothing in the update equations forces DPAE to rise—self-precision is high, and DeGroot/regressor/GNN baselines provide independent comparisons that can and do behave differently (e.g., DeGroot consensus collapse). Ablations and external baselines further show the reported numbers are empirical run outputs, not tautologies of the inputs. There is no self-citation uniqueness theorem, no ansatz smuggled from the authors’ prior work as a forced law, and no renaming of a known closed result as a new derivation. Designing a model that can accumulate local noise and then observing accumulation under real classroom constraints is standard simulation practice, not circularity under the stated criteria.
Axiom & Free-Parameter Ledger
free parameters (6)
- interaction_rounds_R
- partners_per_round_K
- sigma_self
- alpha_i_social_anxiety_noise
- LLM_trust_weight_omega
- observation_precision_tau_from_omega_and_u
axioms (6)
- domain assumption Beliefs about peer ability are well-approximated by independent univariate Gaussians B_jk = N(μ_jk, σ_jk).
- standard math Message influence is precision-weighted Bayesian fusion of Gaussian prior and scalar observation.
- domain assumption Questionnaire friendships and psychometric scales suffice to build individualized subjective graphs G_i and reachability weights π_i.
- ad hoc to paper LLM narrative assessments under RAG locality are valid stand-ins for student communication and credibility judgment.
- domain assumption Group-perceived ranking from averaged belief means is the right collective observable for ‘social perception’ of academic standing.
- domain assumption Exam score map h(Y) places ability on a shared scale usable for cross-student belief fusion.
invented entities (3)
-
Individualized subjective social graph G_i = (V, E_i, π_i)
no independent evidence
-
Rashomon Set Agents / subjective-graph LLM agent population
no independent evidence
-
LLM trust-gating weight ω and induced observation precision τ
no independent evidence
read the original abstract
Social actors do not observe a common social world: each individual forms judgments from a partial and potentially distorted view of the surrounding network. We study whether graph-local evidence and credibility-weighted communication can generate persistent distortions in perceived academic standing, even when agents repeatedly receive objective performance signals. We introduce a data-constrained multi-agent framework in which LLM agents operate through individualized subjective graphs that determine peer visibility, evidence access, and interaction opportunities. Agents exchange uncertainty-annotated assessments, evaluate message credibility, and maintain explicit Gaussian belief states updated through Bayesian fusion. We evaluate the framework on 12 middle-school classrooms comprising 482 students, using questionnaire-derived social information and six consecutive examinations. On the Social-Observed subset (n=419), collective ranking error increases from 0.066 \pm 0.008 to 0.124 \pm 0.009 across six epochs despite repeated exam-based anchoring. Ablations associate individualized visibility and LLM-based trust gating with more stable long-horizon behavior, while constrained retrieval primarily safeguards against global-information leakage. Compared with evaluated DeGroot configurations, the proposed framework achieves lower final ranking error; those DeGroot configurations exhibit near-zero terminal opinion diversity. These findings establish subjective-graph LLM agents as a mechanism-oriented framework for data-constrained simulated social perception. Code is available at https://anonymous.4open.science/r/Rashomonomon-0126.
Figures
Reference graph
Works this paper leans on
-
[1]
Educational data mining to predict bachelors students’ success,
D. Jacob and R. Henriques, “Educational data mining to predict bachelors students’ success,” Emerging Science Journal , vol. 7, pp. 159–171, 2023
2023
-
[2]
Systematic review of research on artificial intelligence applica- tions in higher education—where are the educators?
O. Zawacki-Richter, V. I. Marín, M. Bond, and F. Gouverneur, “Systematic review of research on artificial intelligence applica- tions in higher education—where are the educators?” Interna- tional Journal of Educational Technology in Higher Education , vol. 16, no. 1, p. 39, 2019
2019
-
[3]
It’s who you know: A review of peer networks and academic achievement in schools,
A. Black, M. F. Warstadt, and C. Mamas, “It’s who you know: A review of peer networks and academic achievement in schools,” Frontiers in Psychology, vol. 15, p. 1444570, 2025, doi: 10.3389/fpsyg.2024.1444570
-
[4]
Students’ academic performance prediction using educational data mining and machine learning: A system- atic review,
M. S. N. Al-Din, “Students’ academic performance prediction using educational data mining and machine learning: A system- atic review,” International Journal of Research and Innovation in Social Science , vol. 8, no. 8, pp. 1264–1291, 2024
2024
-
[5]
Multi-agent reinforcement learning: A comprehensive survey,
D. Huh and P. Mohapatra, “Multi-agent reinforcement learning: A comprehensive survey,” arXiv preprint arXiv:2312.10256 , 2023
Pith/arXiv arXiv 2023
-
[6]
L. Vanhée, M. Borit, P.-O. Siebers, R. Cremades, C. Frantz, Ö. Gürcan, F. Kalvas, D. R. Kera, V. Nallur, K. Narasimhan, and M. Neumann, “Large language models for agent-based mod- elling: Current and possible uses across the modelling cycle,” arXiv preprint arXiv:2507.05723 , 2025
Pith/arXiv arXiv 2025
-
[7]
A survey on large language model based autonomous agents,
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin, W. X. Zhao, Z. Wei, and J.-R. Wen, “A survey on large language model based autonomous agents,” Frontiers of Computer Science , vol. 18, no. 6, p. 186345, 2024
2024
-
[8]
Large language model based multi- agents: A survey of progress and challenges,
T. Guo, X. Chen, Y. Wang, R. Chang, S. Pei, N. V. Chawla, O. Wiest, and X. Zhang, “Large language model based multi- agents: A survey of progress and challenges,” in Proc. 33rd Int. Joint Conf. Artificial Intelligence (IJCAI), 2024, pp. 8048–8057
2024
-
[9]
Beyond static responses: Multi-agent LLM systems as a new paradigm for social science research,
J. Haase and S. Pokutta, “Beyond static responses: Multi-agent LLM systems as a new paradigm for social science research,” arXiv preprint arXiv:2506.01839 , 2025
Pith/arXiv arXiv 2025
-
[10]
Graph neural networks for social recommendation,
W. Fan, Y. Ma, Q. Li, Y. He, E. Zhao, J. Tang, and D. Yin, “Graph neural networks for social recommendation,” in Proc. World Wide Web Conf. (WWW) , 2019, pp. 417–426
2019
-
[11]
INFLECT-DGNN: Influencer pre- diction with dynamic graph neural networks,
E. Tiukhova, E. Penaloza, M. Óskarsdóttir, B. Baesens, M. Snoeck, and C. Bravo, “INFLECT-DGNN: Influencer pre- diction with dynamic graph neural networks,” IEEE Access , vol. 12, pp. 115 026–115 041, 2024
2024
-
[12]
AI-based conversational agents: A scoping review from technologies to future directions,
S. Kusal, S. Patil, J. Choudrie, K. Kotecha, S. Mishra, and A. Abraham, “AI-based conversational agents: A scoping review from technologies to future directions,” IEEE Access, vol. 10, pp. 92 337–92 356, 2022
2022
-
[13]
LLMs and generative agent-based models for complex systems research,
Y. Lu, A. Aleta, C. Du, L. Shi, and Y. Moreno, “LLMs and generative agent-based models for complex systems research,” Physics of Life Reviews , vol. 51, pp. 283–293, 2024
2024
-
[14]
Population-aligned persona generation for LLM-based social simulation,
Z. Hu, J. Lian, Z. Xiao, M. Xiong, Y. Lei, T. Wang, K. Ding, Z. Xiao, N. J. Yuan, and X. Xie, “Population-aligned persona generation for LLM-based social simulation,” arXiv preprint arXiv:2509.10127, 2025
arXiv 2025
-
[15]
Reaching a consensus,
M. H. DeGroot, “Reaching a consensus,” Journal of the Ameri- can Statistical Association, vol. 69, no. 345, pp. 118–121, 1974
1974
-
[16]
Social networks in structural equation models,
N. E. Friedkin, “Social networks in structural equation models,” Social Psychology Quarterly , vol. 53, no. 4, pp. 316–328, 1990, doi: 10.2307/2786737
-
[17]
Continuous opinion dynamics under bounded confi- dence: A survey,
J. Lorenz, “Continuous opinion dynamics under bounded confi- dence: A survey,” International Journal of Modern Physics C , vol. 18, no. 12, pp. 1819–1838, 2007
2007
-
[18]
Naive learning in social networks and the wisdom of crowds,
B. Golub and M. O. Jackson, “Naive learning in social networks and the wisdom of crowds,” American Economic Journal: Mi- croeconomics, vol. 2, no. 1, pp. 112–149, 2010
2010
-
[19]
Agent-based modelling meets generative AI in social network simulations,
A. Ferraro, A. Galli, V. La Gatta, M. Postiglione, G. M. Orlando, D. Russo, G. Riccio, A. Romano, and V. Moscato, “Agent-based modelling meets generative AI in social network simulations,” arXiv preprint arXiv:2411.16031 , 2024
Pith/arXiv arXiv 2024
-
[20]
Spread of (mis)information in social networks,
D. Acemoglu, A. Ozdaglar, and A. ParandehGheibi, “Spread of (mis)information in social networks,” Games and Economic Behavior, vol. 70, no. 2, pp. 194–227, 2011
2011
-
[21]
Model- ing interpersonal perception in dyadic interactions: Towards robot-assisted social mediation in the real world,
H. Javed, W. Wang, A. B. Usman, and N. Jamali, “Model- ing interpersonal perception in dyadic interactions: Towards robot-assisted social mediation in the real world,” Frontiers in Robotics and AI , vol. 11, p. 1410957, 2024
2024
-
[22]
Neural relational inference for interacting systems,
T. N. Kipf, E. Fetaya, K.-C. Wang, M. Welling, and R. Zemel, “Neural relational inference for interacting systems,” in Proc. 35th Int. Conf. Machine Learning (ICML) , 2018
2018
-
[23]
Relational inductive biases, deep learning, and graph networks,
P. W. Battaglia, J. B. Hamrick, V. Bapst, A. Sanchez-Gonzalez, V. Zambaldi, M. Malinowski, A. Tacchetti, D. Raposo, A. San- toro, R. Faulkner, C. Gulcehre, F. Song, A. Ballard, J. Gilmer, G. E. Dahl, A. Vaswani, K. Allen, C. Nash, V. Langston, C. Dyer, N. Heess, D. Wierstra, P. Kohli, M. Botvinick, O. Vinyals, L. Li, and R. Pascanu, “Relational inductive ...
Pith/arXiv arXiv 2018
-
[24]
GNN-IR: Examining graph neural networks for influencer recommendations in social media marketing,
J. Park, H. Ahn, D. Kim, and E. Park, “GNN-IR: Examining graph neural networks for influencer recommendations in social media marketing,” Journal of Retailing and Consumer Services , vol. 78, p. 103759, 2024
2024
-
[25]
Generative agents: Interactive simulacra of human behavior,
J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein, “Generative agents: Interactive simulacra of human behavior,” in Proc. 36th Annu. ACM Symp. User Interface Software Technol. (UIST) , 2023
2023
-
[26]
Bayesian teaching enables prob- abilistic reasoning in large language models,
L. Qiu, F. Sha, and K. Allen, “Bayesian teaching enables prob- abilistic reasoning in large language models,” Nature Commu- nications, 2026
2026
-
[27]
H. Huang, X. Shen, S. Wang, L. Meng, D. Liu, D. A. Duchene, H. Wang, and S. Bhatt, “BayesAgent: Bayesian agentic rea- soning under uncertainty via verbalized probabilistic graphical modeling,” arXiv preprint arXiv:2406.05516 , 2024
arXiv 2024
-
[28]
Are LLM belief updates consistent with Bayes’ theorem?
S. Imran, I. Kendiukhov, M. Broerman, A. Thomas, R. Cam- panella, R. Lamb, and P. M. Atkinson, “Are LLM belief updates consistent with Bayes’ theorem?” in Proc. ICML Workshop on World Models, 2025
2025
-
[29]
Uncer- tainty propagation on LLM agent,
Q. Zhao, D. Li, Y. Liu, W. Cheng, Y. Sun, M. Oishi, T. Osaki, K. Matsuda, H. Yao, C. Zhao, H. Chen, and X. Zhao, “Uncer- tainty propagation on LLM agent,” in Proc. 63rd Annu. Meeting Assoc. Comput. Linguist. (ACL) , 2025, pp. 6064–6073
2025
-
[30]
The emerging field of emotion regulation: An integrative review,
J. J. Gross, “The emerging field of emotion regulation: An integrative review,” Review of General Psychology , vol. 2, no. 3, pp. 271–299, 1998
1998
-
[31]
R. W. Picard, Affective Computing . Cambridge, MA, USA: MIT Press, 1997
1997
-
[32]
Large language models show amplified cognitive biases in moral decision-making,
V. Cheung, M. Maier, and F. Lieder, “Large language models show amplified cognitive biases in moral decision-making,” Pro- ceedings of the National Academy of Sciences , vol. 122, no. 22, p. e2412015122, 2025
2025
-
[33]
Mitigating cognitive biases in clinical decision-making through multi-agent conversa- tions using large language models: Simulation study,
Y. Ke, R. Yang, S. A. Lie, T. X. Y. Lim, Y. Ning, I. Li, H. R. Abdullah, D. S. W. Ting, and N. Liu, “Mitigating cognitive biases in clinical decision-making through multi-agent conversa- tions using large language models: Simulation study,” Journal of Medical Internet Research , vol. 26, p. e59439, 2024
2024
-
[34]
LLM agent-based simulation of student activities and mental health using smartphone sensing data,
W. Sommuang, K. Kerdthaisong, P. Buakhaw, A. B. Wong, and N. Yongsatianchot, “LLM agent-based simulation of student activities and mental health using smartphone sensing data,” arXiv preprint arXiv:2508.02679 , 2025
Pith/arXiv arXiv 2025
-
[35]
Explainable multi-agent reinforcement learning for temporal queries,
K. Boggess, S. Kraus, and L. Feng, “Explainable multi-agent reinforcement learning for temporal queries,” in Proc. 32nd Int. Joint Conf. Artificial Intelligence (IJCAI) , 2023, pp. 55–63
2023
-
[36]
An empirical evaluation of the rashomon effect in ex- plainable machine learning,
S. Müller, “An empirical evaluation of the rashomon effect in ex- plainable machine learning,” arXiv preprint arXiv:2306.15786 , 2023
Pith/arXiv arXiv 2023
-
[37]
Cognitive social structures,
D. Krackhardt, “Cognitive social structures,” Social Networks , vol. 9, no. 2, pp. 109–134, 1987
1987
-
[38]
Auto-RAG: Autonomous retrieval-augmented generation for large language models,
T. Yu, S. Zhang, and Y. Feng, “Auto-RAG: Autonomous retrieval-augmented generation for large language models,” arXiv preprint arXiv:2411.19443 , 2024
Pith/arXiv arXiv 2024
-
[39]
Retrieval-augmented gen- eration for large language models: A survey,
Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, M. Wang, and H. Wang, “Retrieval-augmented gen- eration for large language models: A survey,” arXiv preprint arXiv:2312.10997, 2023
Pith/arXiv arXiv 2023
-
[40]
Attention knows whom to trust: Attention-based trust management for LLM multi-agent systems,
P. He, Z. Dai, X. Tang, Y. Xing, H. Liu, J. Zeng, Q. Peng, S. Agrawal, S. Varshney, S. Wang, J. Tang, and Q. He, “Attention knows whom to trust: Attention-based trust management for LLM multi-agent systems,” arXiv preprint arXiv:2506.02546, 2025
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.