Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:19.995889Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 2 inbound Pith citation observations for arXiv:2506.18710.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:19.995889Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T04:17:25.648174Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-10T22:20:49.429444Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 24bb187b-f29c-4585-9952-440af2095de3 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cc8a224a-e939-4fd1-b83f-b674ead897e9 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Lessons from the Trenches on Reproducible Evaluation of Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb00ef13-f91c-4a82-ac67-dd4dde68a7a1 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Distractor generation for multiple-choice questions with predictive prompting and large language models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bea6d80c-6936-4123-94e9-6be154b57933 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Sastry, A
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1624f01f-6c72-46a0-89fb-34ca3ca107d9 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Chang, X
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c9655ea1-32ec-4989-b546-a732b26c6c9d · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab6aa314-1541-4c9d-94fb-38a70c8ef7fc · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b0f66b31-b1ae-4858-8d25-f3f518ecbbc3 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Evaluating Large Language Models: A Comprehensive Survey
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6d1f93e-e3dd-4009-9000-70e8e9a4973f · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Changing Answer Order Can Decrease MMLU Accuracy
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f814db51-67bd-4305-8466-5d994260db4c · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Measuring Massive Multitask Language Understanding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7977192-8b60-4100-8ec4-7b10630d5d89 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Kasenberg, A
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a825aa2b-cc49-4c4d-bcc5-922451871eec · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models MinorBench: A hand-built benchmark for content-based risks for children
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ebea77e-c4d2-444b-ba25-befa5559ed01 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c44b9237-e4a4-420a-a847-6a0db5c20830 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Macina, N
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f1c3676-d795-4007-a5ca-e4cf5528221d · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ec590f1-abd5-4724-a830-c2b6fc3b6637 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Miller and K
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 66338182-1251-4ef6-9821-3d7cb07ade13 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Nancy, D
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dcd55ea8-78ac-492d-8336-9986fa188871 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8c1cade-0438-4809-ab45-d0ef468f386b · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 79513115-4e9f-4085-9aa8-aa8a4158768d · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f14bc13-72fa-4de4-9d86-c1a26c846bcf · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f8914a9b-4d96-4faf-8417-c8c5e8d0b485 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models LearnLM: Improving Gemini for Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80662d06-63a6-4f2f-bbd1-0c5299a439b6 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Evaluating Gemini in an arena for learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 256efb3c-ade7-419c-8204-c485a7f13395 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Large Language Models are not Fair Evaluators
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2210cf71-da4c-40ff-9f41-bf3ab00f527c · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Large Language Models for Education: A Survey and Outlook
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0bf046d-b888-472b-94a8-13ab792af7b1 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f20ff9-76b7-42cc-85f2-bcabae8501c1 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d81164cd-622c-4e6f-8ba6-db336f662f38 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models A Survey of Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d599a70-0eef-4cc6-9f62-cbc8081f806e · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Zheng, H
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fe32bd1a-bcbb-4f4d-8740-b4e7ef7995f3 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Wheredoyouthinkthereismoreplasticine?
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0846e8bf-2909-4ddd-af7e-d10ca4ffaa55 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models Question
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a8232885-ba20-49c7-8ce3-03c1bca6a1c9 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models The text is as follows: —– paragraph —–
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 99315814-de92-45c3-9951-84e64ec29d9f · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models question
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7278ccc6-e8be-452c-8a1c-6b0307cd8019 · outbound
Benchmarking the Pedagogical Knowledge of Large Language Models question
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b77a933a-aeb1-469a-ac90-b173740e58cf · inbound
Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning Benchmarking the Pedagogical Knowledge of Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation da3fb9dd-9cfb-47b3-acdd-2272bc508f34 · inbound
Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education Benchmarking the Pedagogical Knowledge of Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.