Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T12:28:32.395213Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 33 inbound Pith citation observations for arXiv:2305.12474.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T12:28:32.395213Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:21:25.985017Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
19 of 19 outbound references displayed
External citation measurements
19
pith, observed 2026-08-05T02:28:24.338817Z
Observation 5da42aed-6caf-46c4-a61f-bfa49b6d6a31 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1a8e6288-784b-4c65-83d3-4f31f40cdeb7 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c48315cd-a1d6-454d-9eb1-1eb430b1ae93 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e12e333-9264-4843-8436-8c52a50f0f68 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8ee23b33-4950-482b-9801-a2a255f74204 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7f894159-50dc-4532-88d9-738449ee9e38 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2dc7f7f6-d212-4254-919e-b62a1eb59f3d · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark In order to protect this heritage while also developing tourism activities, measures need to be taken to protect the tourism resources
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fdace3f6-08ac-41dd-ad0a-e5043aa0912e · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4e33bdc9-1f4b-4351-98dc-0e868daf10c2 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark This will enhance the cultural literacy and environmental awareness of the tourists and reduce the damage to the terraces
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 39dada78-c3c9-4183-869c-eca92b621828 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7cd947d7-4ba4-49b6-88d3-2d4f26dc2c62 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark At the same time, these facilities should be planned judiciously to avoid damage to the terraces
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a3ac400a-ca29-4c75-9286-86e27fc02386 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark 完善 景区规划、依法保护生态环境
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d84ba65c-aabf-45ed-ae82-c06c3f386dff · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark 普及旅游文化环境保 护教育,提高游客对旅游资源环境保护 的意识
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 299c4ed3-e7e0-43e1-a9a8-f4b8798faa69 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark 评定该‘生 态博物馆’的环境容量,对人口数量的 容纳程度,限制客流量
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b92bfb5d-8eac-4c8e-86f8-4e7640a18f7c · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark 尽可能保证新建设施与景区景观相 融合
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 113cf6b0-eb19-4257-a3b1-60c8a18062e8 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark improve the planning of the scenic area, protect the ecological environment in accordance with the law
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2c8949de-7100-4f53-a883-3c03366751a6 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark popularize education on the protection of the tourism cultural environment, raise tourists’ awareness of the protection of tourism resources and environment
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 920ba379-6e08-437c-acc4-4c22993a1574 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark assess the environmental capacity of this ‘Ecological Museum’, regulate the carrying capacity in terms of population, limit the flow of visitors
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ea8ba8d1-5dab-498d-99f4-bc6d9ff60ae4 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark ensure new facilities blend harmoniously with the scenic landscape
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 42c25863-cfe0-47a6-94c5-2b7f29e04c1c · inbound
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1cd82ad0-a844-4020-973b-484bf5de6693 · inbound
Yi: Open Foundation Models by 01.AI Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0ec704b1-b80c-474d-926c-be559c777325 · inbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ae6bd359-f7f9-4914-8730-115862034a40 · inbound
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8613c96b-2375-4646-93c4-24db5a62c5c3 · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3f5a99a0-ade0-42fc-a449-04d9131a51f9 · inbound
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 03dd5117-4e25-478c-8bb9-bfd808dda126 · inbound
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 207
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e9f3f8a1-0dd0-48bb-80cc-36c9b19869cf · inbound
TASE: Token Awareness and Structured Evaluation for Multilingual Language Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 310061f6-30b1-4ac3-ba9b-a8151a1c8305 · inbound
Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dac95634-5e48-4fae-897d-b95e50e70bcc · inbound
MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca6715dc-d740-4dac-ab4f-437a436c501e · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 178
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ac50c29-e589-40b4-bb08-54936b5c90b9 · inbound
Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea268b62-2206-4296-8766-bdf9dc750395 · inbound
LLaDA2.0: Scaling Up Diffusion Language Models to 100B Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b598bf54-1b21-4e92-a36c-aca177dcc8e8 · inbound
GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 55e59d7a-4086-49e2-9e72-39ecb8321610 · inbound
Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 35fe28d5-b4bb-45b2-a6f0-ccbf62ff98e3 · inbound
TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 92b3a011-13f8-41b1-980d-0c82dc16e045 · inbound
SFT-GRPO Data Overlap as a Post-Training Hyperparameter for Autoformalization Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eabcf859-5c2b-4db3-8fa3-4864e67a1e31 · inbound
RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025) Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5cdbe6e8-0a71-4fde-86f8-073a2554e6d1 · inbound
HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 534afaf3-715d-452a-b419-7f00bad02d07 · inbound
When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 456f76f0-19c1-4c20-8436-2b23c688c80b · inbound
When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70da74c3-a614-47bb-909c-1164359e15da · inbound
Validity-Calibrated Reasoning Distillation Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4a4a2f6e-99c8-48e3-a917-0bfa2a78aede · inbound
Validity-Calibrated Reasoning Distillation Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d8ca5f37-c0f4-41ab-8928-51e0d3ccc0aa · inbound
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c80e584-942d-45ed-92bb-461def9eac17 · inbound
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5cf0c92e-db2f-428c-8fda-99a551535990 · inbound
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3f92752-2110-4304-a475-22c015ef5c2e · inbound
LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations? Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b7ea2fe-d53d-4601-a505-9f236ce6229f · inbound
Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e29a27f3-24a8-4962-837b-10c619cf0093 · inbound
Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a1ecbb9-838e-406b-acd3-c244cf0f52fc · inbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 753a4f1e-a2c1-4cdb-ae93-c89f93ad31be · inbound
Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5e85653-5570-4a7e-956f-0437e5891064 · inbound
EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feebda8d-256e-41e7-8181-d1053c14fbc5 · inbound
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.