Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:30:34.819197Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2412.10056.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:30:34.819197Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f9f34630-1069-48fd-a18f-d4f3e35f5d54 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1beebf1a-d41d-4f68-ac14-b04e49435939 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d06d8d1f-4f0e-4276-aa29-b4a44dac2429 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebbe3dba-72e6-4765-a642-9988a64a8cb9 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? a is b" fail to learn
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 238b180c-a372-431d-81a7-b42414758e1b · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Applying the rasch model: fundamental measurement in the human sciences, 2007
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6ae448e2-4f2d-43ef-b19d-74d80cf957bd · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Boone and Amity Noltemeyer
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 720f12fb-7884-4046-8d3d-c849f7bb9c8b · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Internlm2 technical report, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df7ad01c-87f9-45a1-9afe-1dd4a27d19dc · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a341b58-3540-4bf4-b95a-95630a3b55c8 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? See What LLMs Cannot Answer: A Self-Challenge Framework for Uncovering LLM Weaknesses
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0680fa2f-fecc-4b3b-af2a-2c72670adfba · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Flash A ttention-2: Faster attention with better parallelism and work partitioning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a118e8d-5923-451d-ae2d-a38037b5c9f7 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Large Language Models of Code Fail at Completing Code with Potential Bugs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 812d37ec-fa00-4e6a-8a5a-ed9a8814733b · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3243e408-c118-4e2f-a9e5-379ce6eab065 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea5a6b41-6cc2-4f67-9e9c-6d29285933cb · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Measuring massive multitask language understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8de0acd-576a-40d2-846b-b7107dce6c1e · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Measuring massive multitask language understanding, January 2021 b
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f9c0d431-1110-40f5-a877-8be71f50a559 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Measuring mathematical problem solving with the math dataset
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb6e4b87-812c-471b-83af-6a2bf5457a75 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b2bea7fb-0b39-46d5-a9a7-5182f0f61443 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? New Ontology and Knowledge Graph for University Curriculum Recommendation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 699aae21-23b2-4c57-9c14-495595825d32 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Large language models and simple, stupid bugs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ab75a756-39e4-42ed-b87e-7a1b909235b4 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? FigureQA : An annotated figure dataset for visual reasoning, February 2018
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation acbf3732-4f48-4231-acbd-b8cf6b609d78 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? A diagram is worth a dozen images, March 2016
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 50d3fac8-5555-4968-b1b9-419893fedc70 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Rasch Measurement: Applications in Quantitative Educational Research, volume 1 of Education
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c2524c55-333c-49c2-81fa-21fcdde0b44c · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Cmmlu: Measuring massive multitask language understanding in chinese, 2023
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56a7f43f-c67d-4643-9704-7c7807fcb0d4 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? CMMLU: measuring massive multitask language understanding in chinese
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 862c2ae7-19d3-41f8-8a04-45789ea16988 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Truthfulqa: Measuring how models mimic human falsehoods
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e502e362-2694-42b3-b138-58a17e22547e · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? MMBench : Is your multi-modal model an all-around player?, August 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ab262c57-8b40-40e5-b69e-1991f1982407 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Learn to explain: Multimodal reasoning via thought chains for science question answering, October 2022 a
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3edc29f7-0156-48a2-8c94-6fb23322fa05 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? IconQA : A new benchmark for abstract diagram understanding and visual language reasoning, July 2022 b
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1f8d3dba-57da-4986-99c8-fe07a68f3bb9 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c248b91-8ab9-4a96-a7f0-663ef60f835a · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? OK-VQA : A visual question answering benchmark requiring external knowledge, September 2019
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b4c773d5-cbb0-400c-8c23-35915940be86 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Mistral Large 2: Designed for Single-Node Inference with Long-Context
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 17a8cc32-8cd0-499f-9d89-3ab18a11bf20 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Training on the Benchmark Is Not All You Need
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc5ec2c7-f3f5-437f-b7e4-f7fd51ea6349 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? GPT-4 Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfc35231-87ab-431a-b617-2423d282763e · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 586c737e-b5d0-44fa-b697-59adf5aaeac8 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Probabilistic models for some intelligence and attainment tests
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84fef0db-e386-495d-86d6-3aad81f94a37 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Winogrande: An adversarial winograd schema challenge at scale
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41d8d7db-ae83-4c73-b371-b63728d92ebb · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef9a2e91-6618-47de-a8f3-eadb8d6e1dd0 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Assessing programming task difficulty for efficient evaluation of large language models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eadb0a71-d0e0-4b1e-922a-da1e1c8867d3 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Internlm: A multilingual language model with progressively enhanced capabilities
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 382f101f-6b49-47e6-8b35-2688d437a679 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Mmlu-pro: A more robust and challenging multi-task language understanding benchmark, 2024
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cddee14a-bbc2-48d8-b72c-7f28dd03e332 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 981a2628-7dfc-43d5-8c5c-7073ee72c713 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Qwen2 Technical Report
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a546928d-c373-40bd-9018-3262832ed61d · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b024bb3d-431b-4fd1-996f-03166bc670d0 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66afb50f-3d1f-46ca-9669-24b7d6baeaa3 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 327bc6d7-4af6-4a84-b3e3-d2b9490da6e7 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Can llm replace stack overflow? a study on robustness and reliability of large language model code generation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e9f9fb6f-e8ce-4b77-8a6a-5e870ea1c6ee · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8208fe8-0b2f-4917-b2fa-3a15216799d2 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Larger and more instructable language models become less reliable
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03664885-0d05-47c9-8ada-3b3656cf283b · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Gaokao-mm: A chinese human-level benchmark for multimodal models evaluation, 2024
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ec80e9f8-c573-412e-893c-6516311357aa · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? @esa (Ref
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b9a984e-8ad1-4bc1-ab2a-e77f2536960a · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a33b7fe-d241-4b37-ac08-1263e9b29ec0 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f213a072-b012-4ecd-8f93-9642034b8ab2 · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? , " * write output.state after.block = add.period write
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6af2e208-bf5f-4412-88c1-cf0aa85a9c7e · outbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? write newline
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.