Pith. sign in
Pith Number

pith:EZJGRR34

pith:2026:EZJGRR3453GZ2PZSTTPMA4OTCM
not attested not anchored not stored refs resolved

Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact

Michael Hardy, Yunsung Kim

LLMs share behavioral biases that align poorly with expert human teaching and can oppose intended student learning outcomes.

arxiv:2603.00883 v2 · 2026-03-01 · cs.LG · cs.AI · cs.CY · stat.AP

Add to your LaTeX paper
\usepackage{pith}
\pithnumber{EZJGRR3453GZ2PZSTTPMA4OTCM}

Prints a linked badge after your title and injects PDF metadata. Compiles on arXiv. Learn more · Embed verified badge

Record completeness

1 Bitcoin timestamp
2 Internet Archive
3 Author claim open · sign in to claim
4 Citations open
5 Replications open
Portable graph bundle live · download bundle · merged state
The bundle contains the canonical record plus signed events. A mirror can host it anywhere and recompute the same current state with the deterministic merge algorithm.

Claims

C1strongest claim

Across all LLMs, inter-model behaviors on disparate tasks correlate higher than they do with expert human behaviors on target tasks. These biases shared across LLMs are poorly aligned with downstream measures of teaching quality and often negatively aligned with the intended impact of student learning outcomes.

C2weakest assumption

That the selected difficult-to-verify teaching and learning tasks for schoolchildren accurately capture the intended impact on student outcomes, and that expert human behaviors provide the appropriate reference standard for measuring alignment.

C3one line summary

Shared biases across LLMs from common pretraining misalign with teaching quality and negatively correlate with intended student learning outcomes, with model ensembles amplifying the misalignment.

References

20 extracted · 20 resolved · 2 Pith anchors

[1] The Rapid Adoption of Generative AI. Anthony J. Bishara and James B. Hittner. 2017. Confi- dence intervals for correlations when data are not nor- mal.Behavior Research Methods, 49(1):294–309. David B 2017 · doi:10.1080/10627197.2017.1309274
[2] Technical report, Center for American Progress, Washington, D.C 2021
[3] Are more llm calls all you need? Towards scaling laws of compound inference systems 2023
[4] ISSN: 2692-8205 Pages: 2025.10.16.679418 Section: New Results 2025 · doi:10.1080/10888691.2018.1537791
[5] All that Glitters 2016 · doi:10.1080/10627197.2012.715019

Cited by

1 paper in Pith

Receipt and verification
First computed 2026-07-29T00:24:36.907512Z
Builder pith-number-builder-2026-05-17-v1
Signature Pith Ed25519 (pith-v1-2026-05) · public key
Schema pith-number/v1.0

Canonical hash

265268c77ceecd9d3f329cdec071d313253c25698f3d103a532c6265f2a23e06

Aliases

arxiv: 2603.00883 · arxiv_version: 2603.00883v2 · doi: 10.48550/arxiv.2603.00883 · pith_short_12: EZJGRR3453GZ · pith_short_16: EZJGRR3453GZ2PZS · pith_short_8: EZJGRR34
Agent API
Verify this Pith Number yourself
curl -sH 'Accept: application/ld+json' https://pith.science/pith/EZJGRR3453GZ2PZSTTPMA4OTCM \
  | jq -c '.canonical_record' \
  | python3 -c "import sys,json,hashlib; b=json.dumps(json.loads(sys.stdin.read()), sort_keys=True, separators=(',',':'), ensure_ascii=False).encode(); print(hashlib.sha256(b).hexdigest())"
# expect: 265268c77ceecd9d3f329cdec071d313253c25698f3d103a532c6265f2a23e06
Canonical record JSON
{
  "metadata": {
    "abstract_canon_sha256": "59ee23dd754ba39af55e66b8aa2f70305497299d5b6bea39aacdccf5d1a5d55a",
    "cross_cats_sorted": [
      "cs.AI",
      "cs.CY",
      "stat.AP"
    ],
    "license": "http://creativecommons.org/licenses/by-sa/4.0/",
    "primary_cat": "cs.LG",
    "submitted_at": "2026-03-01T03:05:46Z",
    "title_canon_sha256": "cc353d10cc99d84cfb4a36b58b492b9a199f97e8c138a97b984044181d284555"
  },
  "schema_version": "1.0",
  "source": {
    "id": "2603.00883",
    "kind": "arxiv",
    "version": 2
  }
}