Pith. sign in

Paper Citation Record · LEDGER

Benchmarking the Pedagogical Knowledge of Large Language Models

As of 11 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 2 inbound Pith citation observations for arXiv:2506.18710.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18710 v3

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:19.995889Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:17:25.648174Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T22:20:49.429444Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24bb187b-f29c-4585-9952-440af2095de3 · outbound

This paper cites an unresolved cited work.

Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:20:23.826962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:20:17.479564Z digest=sha256:46ace950c17a0650cabc73e1aea2fd375cbcc12476171625fa2fa5208e19793d

Observation cc8a224a-e939-4fd1-b83f-b674ead897e9 · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

Benchmarking the Pedagogical Knowledge of Large Language Models Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:17.558649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:17.558649Z digest=sha256:8c5741872f7ae3aabd166817950c6b0f5504d7193a63a4c8fc838a624a9abde6

Observation cb00ef13-f91c-4a82-ac67-dd4dde68a7a1 · outbound

This paper cites Distractor generation for multiple-choice questions with predictive prompting and large language models.

Benchmarking the Pedagogical Knowledge of Large Language Models Distractor generation for multiple-choice questions with predictive prompting and large language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:17.679991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:17.679991Z digest=sha256:1d1d3c1447dc7e15e95a9ddc838237019e4dbea2f91c1985e1fc0d947687c2bd

Observation bea6d80c-6936-4123-94e9-6be154b57933 · outbound

This paper cites Sastry, A.

Benchmarking the Pedagogical Knowledge of Large Language Models Sastry, A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:23.535094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:20:17.780784Z digest=sha256:88f9a1e3a83c7360d05dddc6f681bc19d9d5957b0f8a4e2b58b7b798350e12bd

Observation 1624f01f-6c72-46a0-89fb-34ca3ca107d9 · outbound

This paper cites Chang, X.

Benchmarking the Pedagogical Knowledge of Large Language Models Chang, X

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:23.277396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:20:17.859253Z digest=sha256:bae0559525a034098b0ec27185554287c54b4bfdbc582972c41df2b239cb50c6

Observation c9655ea1-32ec-4989-b546-a732b26c6c9d · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Benchmarking the Pedagogical Knowledge of Large Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:17.947382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:17.947382Z digest=sha256:fbe563a4e313ed8991b88dc70ebc2ae85bb89a67f026cf7e6d42507751fd5235

Observation ab6aa314-1541-4c9d-94fb-38a70c8ef7fc · outbound

This paper cites an unresolved cited work.

Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:20:23.093985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:20:18.025882Z digest=sha256:3cb5d08bf4596226efaddc1a888cb04133230554883456a640eade54e8ba4be1

Observation b0f66b31-b1ae-4858-8d25-f3f518ecbbc3 · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

Benchmarking the Pedagogical Knowledge of Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.098093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.098093Z digest=sha256:d036211e24c269e677e73d888e740f2654a009f3c85aa45156fabcd0a4391112

Observation b6d1f93e-e3dd-4009-9000-70e8e9a4973f · outbound

This paper cites Changing Answer Order Can Decrease MMLU Accuracy.

Benchmarking the Pedagogical Knowledge of Large Language Models Changing Answer Order Can Decrease MMLU Accuracy

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.169890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.169890Z digest=sha256:0c3f64f82976d2796ac2298a4c510553b63e10be4ac75c49f80e1f8a16a17038

Observation f814db51-67bd-4305-8466-5d994260db4c · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Benchmarking the Pedagogical Knowledge of Large Language Models Measuring Massive Multitask Language Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.291469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.291469Z digest=sha256:7d71e8e8eb34ccc17cc5a8d134e30b69de874005d4a23faa65e146e16bc62345

Observation a7977192-8b60-4100-8ec4-7b10630d5d89 · outbound

This paper cites Kasenberg, A.

Benchmarking the Pedagogical Knowledge of Large Language Models Kasenberg, A

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.383691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.383691Z digest=sha256:c1966dc532afff189c5ff6480fbd6f6369ef0f067eea05443719e25047ee2f32

Observation a825aa2b-cc49-4c4d-bcc5-922451871eec · outbound

This paper cites MinorBench: A hand-built benchmark for content-based risks for children.

Benchmarking the Pedagogical Knowledge of Large Language Models MinorBench: A hand-built benchmark for content-based risks for children

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.494890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.494890Z digest=sha256:f774b73e1a1c06a1dc9d9315b1ef200d48aef1a22a940d43e5dd851bb9449b15

Observation 5ebea77e-c4d2-444b-ba25-befa5559ed01 · outbound

This paper cites an unresolved cited work.

Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:20:22.906863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:20:18.550027Z digest=sha256:a991089e0d57a0a890042c53855ec529a9dd9b9cffb99b95af902460ce3590d7

Observation c44b9237-e4a4-420a-a847-6a0db5c20830 · outbound

This paper cites Macina, N.

Benchmarking the Pedagogical Knowledge of Large Language Models Macina, N

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.624386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.624386Z digest=sha256:3b35f9336233ebe9ec1ecc42aedd0efad5bca4dc99d0ceb5a5fa83486f1c4889

Observation 2f1c3676-d795-4007-a5ca-e4cf5528221d · outbound

This paper cites Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.

Benchmarking the Pedagogical Knowledge of Large Language Models Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.697366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.697366Z digest=sha256:dc6f967d5bd2cd9bb82d1e6289fbb60736d1cb4dd1985a6de4a776c98ed7475c

Observation 2ec590f1-abd5-4724-a830-c2b6fc3b6637 · outbound

This paper cites Miller and K.

Benchmarking the Pedagogical Knowledge of Large Language Models Miller and K

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:22.542405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:20:18.788322Z digest=sha256:dea62995b1df75d8a857d74548c40074359a1e2818f74a0a9c68c1462d83e4f8

Observation 66338182-1251-4ef6-9821-3d7cb07ade13 · outbound

This paper cites Nancy, D.

Benchmarking the Pedagogical Knowledge of Large Language Models Nancy, D

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:22.232468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:20:18.858576Z digest=sha256:e34e218a617a66596012fa64be8fad88ade63e825b6a481b6b06e46b6c6a0ef8

Observation dcd55ea8-78ac-492d-8336-9986fa188871 · outbound

This paper cites Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions.

Benchmarking the Pedagogical Knowledge of Large Language Models Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.894389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.894389Z digest=sha256:856a40316abc678282b47ac3a6441c1b4f4dcafe7c8d68de0a5021aac7bfb3db

Observation f8c1cade-0438-4809-ab45-d0ef468f386b · outbound

This paper cites an unresolved cited work.

Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:20:21.989970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:20:18.941933Z digest=sha256:df3a2587694c0da13a8d71ad7df605744632d1ac06e3fb97bed49221fbd5010d

Observation 79513115-4e9f-4085-9aa8-aa8a4158768d · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Benchmarking the Pedagogical Knowledge of Large Language Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.015163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.015163Z digest=sha256:19c837c83d233da5ecd3a85a46371ab2cb78a922a39c7128375f2aa2c4df20b9

Observation 2f14bc13-72fa-4de4-9d86-c1a26c846bcf · outbound

This paper cites an unresolved cited work.

Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work

Reference 21

Resolution
verified exact
doi, observed 2026-08-06T23:20:20.188179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:20:19.083382Z digest=sha256:534ed5c651d78c429adf52539be7e0ed5848aba2919cda173f84f5700078230e

Observation f8914a9b-4d96-4faf-8417-c8c5e8d0b485 · outbound

This paper cites LearnLM: Improving Gemini for Learning.

Benchmarking the Pedagogical Knowledge of Large Language Models LearnLM: Improving Gemini for Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.149082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.149082Z digest=sha256:344d79758e344c7ae6d6814ef701c019ee5cf456aff0ca35aa3bfe5dc63d833a

Observation 80662d06-63a6-4f2f-bbd1-0c5299a439b6 · outbound

This paper cites Evaluating Gemini in an arena for learning.

Benchmarking the Pedagogical Knowledge of Large Language Models Evaluating Gemini in an arena for learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.262891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.262891Z digest=sha256:9c2e6e78ef1a52ef4d7143d050f7763a464fea34286cb5c8fe22c518eba67b88

Observation 256efb3c-ade7-419c-8204-c485a7f13395 · outbound

This paper cites Large Language Models are not Fair Evaluators.

Benchmarking the Pedagogical Knowledge of Large Language Models Large Language Models are not Fair Evaluators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.346836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.346836Z digest=sha256:7f32e2983c9510942eb71305ffa995193971685da5bb911573ac9f39c47c1567

Observation 2210cf71-da4c-40ff-9f41-bf3ab00f527c · outbound

This paper cites Large Language Models for Education: A Survey and Outlook.

Benchmarking the Pedagogical Knowledge of Large Language Models Large Language Models for Education: A Survey and Outlook

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.392025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.392025Z digest=sha256:0ad9630371f9bc9aac48f588576345bcafc588244303d38c20828af591d0807d

Observation c0bf046d-b888-472b-94a8-13ab792af7b1 · outbound

This paper cites "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models.

Benchmarking the Pedagogical Knowledge of Large Language Models "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.476553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.476553Z digest=sha256:80487dfec73d64ff1943fb5444479f51074e057ee29b9be1c0cb29a955ead2f5

Observation 77f20ff9-76b7-42cc-85f2-bcabae8501c1 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Benchmarking the Pedagogical Knowledge of Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.533268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.533268Z digest=sha256:e2642b2eccbf128a98c243cd530ebd7ddbf6fc1563e17d406a0a04e6e4991d30

Observation d81164cd-622c-4e6f-8ba6-db336f662f38 · outbound

This paper cites A Survey of Large Language Models.

Benchmarking the Pedagogical Knowledge of Large Language Models A Survey of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.618219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.618219Z digest=sha256:ad8f837aa1a4ec0735f395ff5153937be36b361cb8e999d61e192f932e48cfbd

Observation 6d599a70-0eef-4cc6-9f62-cbc8081f806e · outbound

This paper cites Zheng, H.

Benchmarking the Pedagogical Knowledge of Large Language Models Zheng, H

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:21.782219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:20:19.682163Z digest=sha256:cde0e86636faeb9b02f64153c2c27c1ef0e1fd1262e686bb2852196bd57f2b52

Observation fe32bd1a-bcbb-4f4d-8740-b4e7ef7995f3 · outbound

This paper cites Wheredoyouthinkthereismoreplasticine?.

Benchmarking the Pedagogical Knowledge of Large Language Models Wheredoyouthinkthereismoreplasticine?

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:21.493176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:20:19.753332Z digest=sha256:d46018b4935d78df51dfe7a2561cc256ba0f8e82be4ed46400353f8efdbbbc21

Observation 0846e8bf-2909-4ddd-af7e-d10ca4ffaa55 · outbound

This paper cites Question.

Benchmarking the Pedagogical Knowledge of Large Language Models Question

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:21.120052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:20:19.816947Z digest=sha256:0f2d032eaee9a18997bb7894fc19e82fd3417859f5b3b71b488b04b4e0ec657b

Observation a8232885-ba20-49c7-8ce3-03c1bca6a1c9 · outbound

This paper cites The text is as follows: —– paragraph —–.

Benchmarking the Pedagogical Knowledge of Large Language Models The text is as follows: —– paragraph —–

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:20.945282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:20:19.868065Z digest=sha256:483ad7cf568d934dfaf993051672fce8ef607509490a46417fee9f2598df512e

Observation 99315814-de92-45c3-9951-84e64ec29d9f · outbound

This paper cites question.

Benchmarking the Pedagogical Knowledge of Large Language Models question

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:20.827125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:20:19.914280Z digest=sha256:07ec89a2c6598c7b41d9e990dcf89a5828c3f0dc4ce65472d409eb2c0ace941c

Observation 7278ccc6-e8be-452c-8a1c-6b0307cd8019 · outbound

This paper cites question.

Benchmarking the Pedagogical Knowledge of Large Language Models question

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:20.698192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:20:19.995889Z digest=sha256:0a856a56dcb0e02f3f6f7608620a3647bcde7034d6e6a78d32fc1fbb94d68cef

Pith citing papers

Observation b77a933a-aeb1-469a-ac90-b173740e58cf · inbound

Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning cites this paper.

Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning Benchmarking the Pedagogical Knowledge of Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:49.432408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:55:17.077281Z digest=sha256:f724a8c45a60fd5401a0a4a7b35b6cc9ad09d8ff0cc8fa13b40411c2d52788b8

Observation da3fb9dd-9cfb-47b3-acdd-2272bc508f34 · inbound

Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education cites this paper.

Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education Benchmarking the Pedagogical Knowledge of Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T04:17:25.648174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:17:25.648174Z digest=sha256:2a5be4d902137a592f498cd40a79b5bffe7fac9b49d1bdce29b0a2a9ead3371a