{"paper":{"title":"Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact","license":"http://creativecommons.org/licenses/by-sa/4.0/","headline":"LLMs share behavioral biases that align poorly with expert human teaching and can oppose intended student learning outcomes.","cross_cats":["cs.AI","cs.CY","stat.AP"],"primary_cat":"cs.LG","authors_text":"Michael Hardy, Yunsung Kim","submitted_at":"2026-03-01T03:05:46Z","abstract_excerpt":"LLMs increasingly excel on AI benchmarks, but doing so does not guarantee validity for downstream tasks. This study contrasts LLM alignment on benchmarks, downstream tasks, and, importantly the intended impact of those tasks. We evaluate the performance of leading LLMs (i.e., generative pre-trained base models) on difficult-to-verify tasks of the teaching and learning of schoolchildren. Across all LLMs, inter-model behaviors on disparate tasks correlate higher than they do with expert human behaviors on target tasks. These biases shared across LLMs are poorly aligned with downstream measures o"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Across all LLMs, inter-model behaviors on disparate tasks correlate higher than they do with expert human behaviors on target tasks. These biases shared across LLMs are poorly aligned with downstream measures of teaching quality and often negatively aligned with the intended impact of student learning outcomes.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the selected difficult-to-verify teaching and learning tasks for schoolchildren accurately capture the intended impact on student outcomes, and that expert human behaviors provide the appropriate reference standard for measuring alignment.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Shared biases across LLMs from common pretraining misalign with teaching quality and negatively correlate with intended student learning outcomes, with model ensembles amplifying the misalignment.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"LLMs share behavioral biases that align poorly with expert human teaching and can oppose intended student learning outcomes.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"5e18d16c9f1a0897b191861ce919c675fd8bf0bb7c703f95860f1de80e2f0773"},"source":{"id":"2603.00883","kind":"arxiv","version":2},"verdict":{"id":"ed0fd2bf-4b69-498b-83ac-6318a32d0e43","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-15T18:41:26.515521Z","strongest_claim":"Across all LLMs, inter-model behaviors on disparate tasks correlate higher than they do with expert human behaviors on target tasks. These biases shared across LLMs are poorly aligned with downstream measures of teaching quality and often negatively aligned with the intended impact of student learning outcomes.","one_line_summary":"Shared biases across LLMs from common pretraining misalign with teaching quality and negatively correlate with intended student learning outcomes, with model ensembles amplifying the misalignment.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the selected difficult-to-verify teaching and learning tasks for schoolchildren accurately capture the intended impact on student outcomes, and that expert human behaviors provide the appropriate reference standard for measuring alignment.","pith_extraction_headline":"LLMs share behavioral biases that align poorly with expert human teaching and can oppose intended student learning outcomes."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2603.00883/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":20,"sample":[{"doi":"10.1080/10627197.2017.1309274","year":2017,"title":"The Rapid Adoption of Generative AI. Anthony J. Bishara and James B. Hittner. 2017. Confi- dence intervals for correlations when data are not nor- mal.Behavior Research Methods, 49(1):294–309. David B","work_id":"8ad7af2e-522a-4419-94aa-e9dc5d32ebd5","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2021,"title":"Technical report, Center for American Progress, Washington, D.C","work_id":"16d76f6a-0a3f-4d0e-84d7-65649a553c20","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2023,"title":"Are more llm calls all you need? Towards scaling laws of compound inference systems","work_id":"11a6e30d-1d15-49ba-9408-0f1386c2c043","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"10.1080/10888691.2018.1537791","year":2025,"title":"ISSN: 2692-8205 Pages: 2025.10.16.679418 Section: New Results","work_id":"6ad5b8f3-d267-468b-a522-ec4d0cda308d","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"10.1080/10627197.2012.715019","year":2016,"title":"All that Glitters","work_id":"347ac587-21f2-4d60-a490-ffdc45ac1fb0","ref_index":5,"cited_arxiv_id":"","is_internal_anchor":false}],"resolved_work":20,"snapshot_sha256":"f559dbc6425d36dd8e3e137ea0ab5f4c09e6f8b7ef5d61a839f938cfd6c46363","internal_anchors":2},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}