Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T12:28:32.395213Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 53 inbound Pith citation observations for arXiv:2305.12474.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T12:28:32.395213Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:13:25.067403Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
19 of 19 outbound references displayed
External citation measurements
19
pith, observed 2026-08-05T02:28:24.338817Z
Observation 5da42aed-6caf-46c4-a61f-bfa49b6d6a31 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1a8e6288-784b-4c65-83d3-4f31f40cdeb7 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c48315cd-a1d6-454d-9eb1-1eb430b1ae93 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2e12e333-9264-4843-8436-8c52a50f0f68 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8ee23b33-4950-482b-9801-a2a255f74204 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7f894159-50dc-4532-88d9-738449ee9e38 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2dc7f7f6-d212-4254-919e-b62a1eb59f3d · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark In order to protect this heritage while also developing tourism activities, measures need to be taken to protect the tourism resources
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fdace3f6-08ac-41dd-ad0a-e5043aa0912e · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4e33bdc9-1f4b-4351-98dc-0e868daf10c2 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark This will enhance the cultural literacy and environmental awareness of the tourists and reduce the damage to the terraces
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 39dada78-c3c9-4183-869c-eca92b621828 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7cd947d7-4ba4-49b6-88d3-2d4f26dc2c62 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark At the same time, these facilities should be planned judiciously to avoid damage to the terraces
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a3ac400a-ca29-4c75-9286-86e27fc02386 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark 完善 景区规划、依法保护生态环境
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d84ba65c-aabf-45ed-ae82-c06c3f386dff · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark 普及旅游文化环境保 护教育,提高游客对旅游资源环境保护 的意识
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 299c4ed3-e7e0-43e1-a9a8-f4b8798faa69 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark 评定该‘生 态博物馆’的环境容量,对人口数量的 容纳程度,限制客流量
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b92bfb5d-8eac-4c8e-86f8-4e7640a18f7c · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark 尽可能保证新建设施与景区景观相 融合
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 113cf6b0-eb19-4257-a3b1-60c8a18062e8 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark improve the planning of the scenic area, protect the ecological environment in accordance with the law
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2c8949de-7100-4f53-a883-3c03366751a6 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark popularize education on the protection of the tourism cultural environment, raise tourists’ awareness of the protection of tourism resources and environment
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 920ba379-6e08-437c-acc4-4c22993a1574 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark assess the environmental capacity of this ‘Ecological Museum’, regulate the carrying capacity in terms of population, limit the flow of visitors
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ea8ba8d1-5dab-498d-99f4-bc6d9ff60ae4 · outbound
Evaluating the Performance of Large Language Models on GAOKAO Benchmark ensure new facilities blend harmoniously with the scenic landscape
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 42c25863-cfe0-47a6-94c5-2b7f29e04c1c · inbound
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1cd82ad0-a844-4020-973b-484bf5de6693 · inbound
Yi: Open Foundation Models by 01.AI Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0ec704b1-b80c-474d-926c-be559c777325 · inbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ae6bd359-f7f9-4914-8730-115862034a40 · inbound
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dc451ee2-1095-4b4b-9517-6ec65f0781e4 · inbound
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66afb50f-3d1f-46ca-9669-24b7d6baeaa3 · inbound
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 638e21c2-aa9c-4581-8b88-c5f62bd91e0b · inbound
CDS: Knowledge Component-Driven Data Synthesis Guided by Cognitive Diagnosis Theory Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4854b141-2be6-49bd-86a6-f6021cf02d0e · inbound
O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fe7a5e4-84dc-4c3a-9b96-03c0bc0a401a · inbound
UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a209798-5e44-4c3b-b160-74f6c9fec81c · inbound
Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8613c96b-2375-4646-93c4-24db5a62c5c3 · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b7cc6d3b-11e8-4fe4-8f46-1fd3b437228e · inbound
QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c269c917-9381-49f4-8d96-57aa088222ea · inbound
SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32f788e3-567f-4c5a-9aa6-e71b46c87d7f · inbound
Can reasoning models comprehend mathematical problems in Chinese ancient texts? An empirical study based on data from Suanjing Shishu Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d5e8772-f66c-456f-a66f-85d43e3b6f5e · inbound
Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bcc08b4-3cd8-47fb-9fa7-05f0844bd60c · inbound
Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95bdd5df-221f-4d22-8a4a-88ac65c41668 · inbound
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8acb2a4-80f8-41d0-b9af-dd3134dfd0cc · inbound
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4df059e3-555e-4593-9baf-d468eea40bea · inbound
From Objectives to Questions: A Planning-based Framework for Educational Mathematical Question Generation Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c5974ad-f0e9-423e-8360-06266f88d0fe · inbound
Temporalizing Confidence: Evaluation of Chain-of-Thought Reasoning with Signal Temporal Logic Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aad07c4-6676-4390-ac88-39428790589a · inbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f5a99a0-ade0-42fc-a449-04d9131a51f9 · inbound
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5e9d4b98-9a5e-4136-ba7c-94d16b4bbfc6 · inbound
MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b86644a4-87fc-48f9-8936-06d9325cf7b7 · inbound
Enterprise Large Language Model Evaluation Benchmark Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03dd5117-4e25-478c-8bb9-bfd808dda126 · inbound
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 207
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9435e165-90f4-4332-b5e9-359316fbaa19 · inbound
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc64b61d-def9-4836-8807-fea7c042652c · inbound
Technical Report of TeleChat2, TeleChat2.5 and T1 Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9f3f8a1-0dd0-48bb-80cc-36c9b19869cf · inbound
TASE: Token Awareness and Structured Evaluation for Multilingual Language Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 310061f6-30b1-4ac3-ba9b-a8151a1c8305 · inbound
Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dac95634-5e48-4fae-897d-b95e50e70bcc · inbound
MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca6715dc-d740-4dac-ab4f-437a436c501e · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 178
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6ac50c29-e589-40b4-bb08-54936b5c90b9 · inbound
Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea268b62-2206-4296-8766-bdf9dc750395 · inbound
LLaDA2.0: Scaling Up Diffusion Language Models to 100B Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b598bf54-1b21-4e92-a36c-aca177dcc8e8 · inbound
GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 55e59d7a-4086-49e2-9e72-39ecb8321610 · inbound
Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 35fe28d5-b4bb-45b2-a6f0-ccbf62ff98e3 · inbound
TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 92b3a011-13f8-41b1-980d-0c82dc16e045 · inbound
SFT-GRPO Data Overlap as a Post-Training Hyperparameter for Autoformalization Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation eabcf859-5c2b-4db3-8fa3-4864e67a1e31 · inbound
RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025) Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5cdbe6e8-0a71-4fde-86f8-073a2554e6d1 · inbound
HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 534afaf3-715d-452a-b419-7f00bad02d07 · inbound
When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 456f76f0-19c1-4c20-8436-2b23c688c80b · inbound
When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70da74c3-a614-47bb-909c-1164359e15da · inbound
Validity-Calibrated Reasoning Distillation Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4a4a2f6e-99c8-48e3-a917-0bfa2a78aede · inbound
Validity-Calibrated Reasoning Distillation Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d8ca5f37-c0f4-41ab-8928-51e0d3ccc0aa · inbound
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0c80e584-942d-45ed-92bb-461def9eac17 · inbound
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5cf0c92e-db2f-428c-8fda-99a551535990 · inbound
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3f92752-2110-4304-a475-22c015ef5c2e · inbound
LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations? Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1b7ea2fe-d53d-4601-a505-9f236ce6229f · inbound
Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e29a27f3-24a8-4962-837b-10c619cf0093 · inbound
Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9a1ecbb9-838e-406b-acd3-c244cf0f52fc · inbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 753a4f1e-a2c1-4cdb-ae93-c89f93ad31be · inbound
Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5e85653-5570-4a7e-956f-0437e5891064 · inbound
EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feebda8d-256e-41e7-8181-d1053c14fbc5 · inbound
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.