Pith. sign in

Paper Citation Record · LEDGER

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

As of 23 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 53 inbound Pith citation observations for arXiv:2305.12474.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.12474 v3

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T12:28:32.395213Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 53 of 53 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:13:25.067403Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

19
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5da42aed-6caf-46c4-a61f-bfa49b6d6a31 · outbound

This paper cites Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.495884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:5b6cca502e6878feee8df9cfc1a96d37b8c511a7cd4de030a82aaf42bcbd24d9

Observation 1a8e6288-784b-4c65-83d3-4f31f40cdeb7 · outbound

This paper cites GPT-4 Technical Report.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark GPT-4 Technical Report

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T12:28:32.420686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:a2503fdae31c832777183d83a1f8a4e441d2eeb760c30e351ece67a78085f99c

Observation c48315cd-a1d6-454d-9eb1-1eb430b1ae93 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T12:28:32.430613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:15920f6b296c76ec1cb82cdc1914a581c6baf23956ab5c2b6ee1d30086385d14

Observation 2e12e333-9264-4843-8436-8c52a50f0f68 · outbound

This paper cites an unresolved cited work.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-17T12:28:32.507446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:c5137d5bc5d4c6566e7d5f1ae3443862250191b4ec6a17897a82e531b83c61cb

Observation 8ee23b33-4950-482b-9801-a2a255f74204 · outbound

This paper cites an unresolved cited work.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-17T12:28:32.437185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:15d77fa5481f2ac4e77a4c0933ba6208a76f88c8be38ac21587524241aa401e4

Observation 7f894159-50dc-4532-88d9-738449ee9e38 · outbound

This paper cites an unresolved cited work.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-17T12:28:32.441022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:5b03c1b65bbbaa2a6f8f53ae7f53999429436afe42965064ff8167e85b87b568

Observation 2dc7f7f6-d212-4254-919e-b62a1eb59f3d · outbound

This paper cites In order to protect this heritage while also developing tourism activities, measures need to be taken to protect the tourism resources.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark In order to protect this heritage while also developing tourism activities, measures need to be taken to protect the tourism resources

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.446123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:18276d12c6b8c10ff3804511b367daeefe7f383f28751fcc91695c09d7d53018

Observation fdace3f6-08ac-41dd-ad0a-e5043aa0912e · outbound

This paper cites an unresolved cited work.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-17T12:28:32.449944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:37c8be7e50f6eb9ed623bea3259b8d0686db8cc57a1d8d5e204e873d5611fbd0

Observation 4e33bdc9-1f4b-4351-98dc-0e868daf10c2 · outbound

This paper cites This will enhance the cultural literacy and environmental awareness of the tourists and reduce the damage to the terraces.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark This will enhance the cultural literacy and environmental awareness of the tourists and reduce the damage to the terraces

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.453796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:c14541eda92353787fd2a0f4e88eec6d3d5daa16e75fa3c36d80a06cb9f285b7

Observation 39dada78-c3c9-4183-869c-eca92b621828 · outbound

This paper cites an unresolved cited work.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-17T12:28:32.457230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:ceb7963cad89f0fc964a2c2fe23d09001bcebed2e470725984deb9ea840d5f1c

Observation 7cd947d7-4ba4-49b6-88d3-2d4f26dc2c62 · outbound

This paper cites At the same time, these facilities should be planned judiciously to avoid damage to the terraces.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark At the same time, these facilities should be planned judiciously to avoid damage to the terraces

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.462517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:4a767a72447f53338cdb03e591284921d1380ffc5bd43bfcdcffd30b00a7977f

Observation a3ac400a-ca29-4c75-9286-86e27fc02386 · outbound

This paper cites 完善 景区规划、依法保护生态环境.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark 完善 景区规划、依法保护生态环境

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.466212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:f28bea5aff7e9642b7535f746a241b30b57974e889d02c521ec2160aa199fdd1

Observation d84ba65c-aabf-45ed-ae82-c06c3f386dff · outbound

This paper cites 普及旅游文化环境保 护教育,提高游客对旅游资源环境保护 的意识.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark 普及旅游文化环境保 护教育,提高游客对旅游资源环境保护 的意识

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.469890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:44ffa1e92f2cd39e6c976e301900accf521a9241c9bbd89ef6213916a16535e6

Observation 299c4ed3-e7e0-43e1-a9a8-f4b8798faa69 · outbound

This paper cites 评定该‘生 态博物馆’的环境容量,对人口数量的 容纳程度,限制客流量.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark 评定该‘生 态博物馆’的环境容量,对人口数量的 容纳程度,限制客流量

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.473506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:57136a7a5ae211d6d7265a88c2af547d9935c9c3d3ae4984a2ac9189da8f9fd7

Observation b92bfb5d-8eac-4c8e-86f8-4e7640a18f7c · outbound

This paper cites 尽可能保证新建设施与景区景观相 融合.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark 尽可能保证新建设施与景区景观相 融合

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.478599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:8447954154eeea46e5632d769bbf182271de18aa568218ba229d6ef1e328c562

Observation 113cf6b0-eb19-4257-a3b1-60c8a18062e8 · outbound

This paper cites improve the planning of the scenic area, protect the ecological environment in accordance with the law.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark improve the planning of the scenic area, protect the ecological environment in accordance with the law

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.482604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:d81be412e8c43a35589e122d0f5e297fcfa71ca2d9567903e551e44b3371d4df

Observation 2c8949de-7100-4f53-a883-3c03366751a6 · outbound

This paper cites popularize education on the protection of the tourism cultural environment, raise tourists’ awareness of the protection of tourism resources and environment.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark popularize education on the protection of the tourism cultural environment, raise tourists’ awareness of the protection of tourism resources and environment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.487677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:0dc31828668958a38a5aa0b2b87dd11f9b27d491a6b5cbf174b04ce0681a6ae8

Observation 920ba379-6e08-437c-acc4-4c22993a1574 · outbound

This paper cites assess the environmental capacity of this ‘Ecological Museum’, regulate the carrying capacity in terms of population, limit the flow of visitors.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark assess the environmental capacity of this ‘Ecological Museum’, regulate the carrying capacity in terms of population, limit the flow of visitors

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.491874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:7bfe79fb93506fd41a4a5efbe02575f55ee4bbe15d671301b27e4eeb2a56a7d0

Observation ea8ba8d1-5dab-498d-99f4-bc6d9ff60ae4 · outbound

This paper cites ensure new facilities blend harmoniously with the scenic landscape.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark ensure new facilities blend harmoniously with the scenic landscape

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.503672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:a8b97725b55e5a622b9138515fb132f1819cdf1e95223d10a9f05f04094b014e

Pith citing papers

Observation 42c25863-cfe0-47a6-94c5-2b7f29e04c1c · inbound

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism cites this paper.

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T06:08:05.550346Z digest=sha256:71cf13d0954dad6939a80f114dae69c480f6e8e7800ead133d218ac9faf13780

Observation 1cd82ad0-a844-4020-973b-484bf5de6693 · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:26b9af0d550158970bca4ea89ebfe0033a84440dac559d31ae68c532289882bf

Observation 0ec704b1-b80c-474d-926c-be559c777325 · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:d5028baa0019f3eda7cfceaddddff7d7069f581e26e4c74797a9e1768c07daca

Observation ae6bd359-f7f9-4914-8730-115862034a40 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:cd6eb3dee1b9ecd283e68dee6b0ae1d7bac01173f6ed74ecacf9f9552cf66a9b

Observation dc451ee2-1095-4b4b-9517-6ec65f0781e4 · inbound

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation cites this paper.

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T22:37:56.514739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:37:56.514739Z digest=sha256:d55fd8cd911b27ebb169684be62d1747ecae4b6cb1218996f582a773beda7a94

Observation 66afb50f-3d1f-46ca-9669-24b7d6baeaa3 · inbound

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? cites this paper.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.768686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.768686Z digest=sha256:8d48dc8d9bace95dec8cb052eaa41341757ff0c65800098de40eb4a492857557

Observation 638e21c2-aa9c-4581-8b88-c5f62bd91e0b · inbound

CDS: Knowledge Component-Driven Data Synthesis Guided by Cognitive Diagnosis Theory cites this paper.

CDS: Knowledge Component-Driven Data Synthesis Guided by Cognitive Diagnosis Theory Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T20:44:17.059904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:44:17.059904Z digest=sha256:d6bcab86f6f634fa09f21883d7c2ae570e7b9a93628c1bd894b789a42efca265

Observation 4854b141-2be6-49bd-86a6-f6021cf02d0e · inbound

O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning cites this paper.

O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T17:09:06.856556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:09:06.856556Z digest=sha256:ad6913b2f37a8560ca5b1d4b7e9d94a5df780a0ca7f6e4284eab65dbdf1b8649

Observation 4fe7a5e4-84dc-4c3a-9b96-03c0bc0a401a · inbound

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models cites this paper.

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-09T19:27:46.909958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:27:46.909958Z digest=sha256:941a1c7af8ed03ee3d986caf1a63568b04f6902b5574518e8144a294ffdb9014

Observation 4a209798-5e44-4c3b-b160-74f6c9fec81c · inbound

Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis cites this paper.

Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-09T00:35:28.543376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T00:35:28.543376Z digest=sha256:10627d386282a6bdb1a5512f0485016e3ee07f0e6a53ca16e52642da869b562f

Observation 8613c96b-2375-4646-93c4-24db5a62c5c3 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:1cb37303cf0c53dcee1a5261249be02c32ae9bcf6a7ecda347591c17e744b68d

Observation b7cc6d3b-11e8-4fe4-8f46-1fd3b437228e · inbound

QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation cites this paper.

QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T23:13:25.067403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:13:25.067403Z digest=sha256:bbd54439879e7f1ef63900a0724d62778aaf7f64bf56407864cde305cba07fc0

Observation c269c917-9381-49f4-8d96-57aa088222ea · inbound

SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models cites this paper.

SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T22:25:37.456671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:25:37.456671Z digest=sha256:96f8346bc91c93045fa2c48805e3c8f65c7df40c1c59c47e396679d346567c99

Observation 32f788e3-567f-4c5a-9aa6-e71b46c87d7f · inbound

Can reasoning models comprehend mathematical problems in Chinese ancient texts? An empirical study based on data from Suanjing Shishu cites this paper.

Can reasoning models comprehend mathematical problems in Chinese ancient texts? An empirical study based on data from Suanjing Shishu Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:06.140735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:06.140735Z digest=sha256:958fbbee4a11b0ec35d5fa4719f20c33fca511e0ff01beac34a80fd122ec8dfd

Observation 1d5e8772-f66c-456f-a66f-85d43e3b6f5e · inbound

Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting cites this paper.

Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:22.750533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:22.750533Z digest=sha256:303b2589ac96c775d5f9b1527711f2542d0787070fd919e94599222fc120dfa2

Observation 9bcc08b4-3cd8-47fb-9fa7-05f0844bd60c · inbound

Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models cites this paper.

Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.638949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.638949Z digest=sha256:e622432000249b44db7f3222ccab0e27a8a12526965d070904b06cef19659ba0

Observation 95bdd5df-221f-4d22-8a4a-88ac65c41668 · inbound

Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning cites this paper.

Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.756125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:02.756125Z digest=sha256:cd200c6fbc76552e4316d0cdde03a5d67feeadc5846bebc2a70ad365edf3117c

Observation e8acb2a4-80f8-41d0-b9af-dd3134dfd0cc · inbound

Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation cites this paper.

Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:01.234873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:01.234873Z digest=sha256:0505ea8fad2ca07167044fb99d04597c4fc6fe65af0237f2167243b902981258

Observation 4df059e3-555e-4593-9baf-d468eea40bea · inbound

From Objectives to Questions: A Planning-based Framework for Educational Mathematical Question Generation cites this paper.

From Objectives to Questions: A Planning-based Framework for Educational Mathematical Question Generation Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:08.869060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:08.869060Z digest=sha256:0b45f55807af8c6cfceb374b96e75c2e128f6436ac99013aa14d1a0686732ef2

Observation 7c5974ad-f0e9-423e-8360-06266f88d0fe · inbound

Temporalizing Confidence: Evaluation of Chain-of-Thought Reasoning with Signal Temporal Logic cites this paper.

Temporalizing Confidence: Evaluation of Chain-of-Thought Reasoning with Signal Temporal Logic Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:04.567584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:19:04.567584Z digest=sha256:61f6e67203175d690fb32382125392aac0375e96b96d38db8332723f5991c4e2

Observation 9aad07c4-6676-4390-ac88-39428790589a · inbound

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning cites this paper.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.287403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.287403Z digest=sha256:a2cc2d29789209a9eaa395c6fc453c630ec677fe155132743552249627a1acc7

Observation 3f5a99a0-ade0-42fc-a449-04d9131a51f9 · inbound

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource cites this paper.

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:05:47.729661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T00:05:08.916339Z digest=sha256:3afd1040feb1deaa9178f1b9472885f524599e5957a5eec4bf2fbfa99e3badb0

Observation 5e9d4b98-9a5e-4136-ba7c-94d16b4bbfc6 · inbound

MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs cites this paper.

MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:00.484031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:00.484031Z digest=sha256:2ebbdff83e5f6affa6362f195f9ad8b2cf3f07753524ce23b435e39e5f8ccfb9

Observation b86644a4-87fc-48f9-8936-06d9325cf7b7 · inbound

Enterprise Large Language Model Evaluation Benchmark cites this paper.

Enterprise Large Language Model Evaluation Benchmark Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:32.401251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:32.401251Z digest=sha256:2aea4020a832aa109b40b12d96e8db0a2b995b9820920fcc1d458c5d94f5ad64

Observation 03dd5117-4e25-478c-8bb9-bfd808dda126 · inbound

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning cites this paper.

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 207

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:01:10.027080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-19T01:01:09.840919Z digest=sha256:025d5fa77a9f977079edec011481b3caea4f1cd379b0cda4dcd22f32f012fb8f

Observation 9435e165-90f4-4332-b5e9-359316fbaa19 · inbound

WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training cites this paper.

WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T14:49:40.371463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:49:40.371463Z digest=sha256:47efc972816aa4260fe088c3c5baa3f58a23533b69d7883c698c1b98f899deae

Observation bc64b61d-def9-4836-8807-fea7c042652c · inbound

Technical Report of TeleChat2, TeleChat2.5 and T1 cites this paper.

Technical Report of TeleChat2, TeleChat2.5 and T1 Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:27.053349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:43:27.053349Z digest=sha256:d710cfba2e77f291fda78d603df35910fe4e373fd333e4903af8eda8f1cd524f

Observation e9f3f8a1-0dd0-48bb-80cc-36c9b19869cf · inbound

TASE: Token Awareness and Structured Evaluation for Multilingual Language Models cites this paper.

TASE: Token Awareness and Structured Evaluation for Multilingual Language Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T23:21:25.985017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:21:25.985017Z digest=sha256:c41c68a3cce742ff430df0323dfd1fa28382a31a297e79706002a96a117c8081

Observation 310061f6-30b1-4ac3-ba9b-a8151a1c8305 · inbound

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support cites this paper.

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-05T23:04:57.082913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:04:57.082913Z digest=sha256:e5eec48d58e637d7cf498f008899a203b06daccee79f76e8d8f4c3d3e651d780

Observation dac95634-5e48-4fae-897d-b95e50e70bcc · inbound

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models cites this paper.

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T18:53:05.214761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:53:05.214761Z digest=sha256:e6b2b76cb5cedb1dca7694ab56ed4d6d59f352b8a08f9c13a4e96bbe3be2c867

Observation ca6715dc-d740-4dac-ab4f-437a436c501e · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:31735239c6abbcd05f5f1ad81e082e828f80adbb1d3408e461fce5958b59f305

Observation 6ac50c29-e589-40b4-bb08-54936b5c90b9 · inbound

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards cites this paper.

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:48.547800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:15:48.547800Z digest=sha256:0b7129b9d48212136c1b8de936da3f8761f5fdc2035ea674ebed1f8f2aa24b0d

Observation ea268b62-2206-4296-8766-bdf9dc750395 · inbound

LLaDA2.0: Scaling Up Diffusion Language Models to 100B cites this paper.

LLaDA2.0: Scaling Up Diffusion Language Models to 100B Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T18:53:20.911374Z digest=sha256:327951d482a274e33414c1f7525ee74b8633e1c55b079617e57af45aec20caae

Observation b598bf54-1b21-4e92-a36c-aca177dcc8e8 · inbound

GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents cites this paper.

GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:05:26.681888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T07:02:28.660002Z digest=sha256:d9ef0237d66abf9295bb2afbe2d97d9a77b3bc3b60eacd5a7587b084e65f7391

Observation 55e59d7a-4086-49e2-9e72-39ecb8321610 · inbound

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding cites this paper.

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T09:11:31.870441Z digest=sha256:d92fe3336d0b616ec187379a28680af84f3e946106382448aaf9ed6b1ab25617

Observation 35fe28d5-b4bb-45b2-a6f0-ccbf62ff98e3 · inbound

TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice cites this paper.

TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T17:59:44.844149Z digest=sha256:8d447ded2a86b621754ff02ecd03e0d05f51b2297604b1287631fbd2751a7231

Observation 92b3a011-13f8-41b1-980d-0c82dc16e045 · inbound

SFT-GRPO Data Overlap as a Post-Training Hyperparameter for Autoformalization cites this paper.

SFT-GRPO Data Overlap as a Post-Training Hyperparameter for Autoformalization Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T13:21:52.225115Z digest=sha256:fefc0b49cbbce54cfafd863691b4c96fa5531102daa9f371e862275de5221a35

Observation eabcf859-5c2b-4db3-8fa3-4864e67a1e31 · inbound

RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025) cites this paper.

RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025) Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 33

Resolution
malformed identifier
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T22:01:21.565986Z digest=sha256:759ed291f33739bce9ddf0cefa30df4b4197380563133ca0a8ce9596c7e647d0

Observation 5cdbe6e8-0a71-4fde-86f8-073a2554e6d1 · inbound

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs cites this paper.

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T00:51:37.506096Z digest=sha256:0169123b42a001dadb386126df1a1c296b9a14f7554c3bb78cd85c5cc4e818b9

Observation 534afaf3-715d-452a-b419-7f00bad02d07 · inbound

When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling cites this paper.

When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T11:00:21.413246Z digest=sha256:662346d367b2deb79c67a0e0390de86c5fc47d0202ded44d95f7383eb0443d2b

Observation 456f76f0-19c1-4c20-8436-2b23c688c80b · inbound

When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling cites this paper.

When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T05:23:31.540079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:23:31.540079Z digest=sha256:74be50c4093786140edb64d447f4643845ede99719679cdcf77df9be91569c46

Observation 70da74c3-a614-47bb-909c-1164359e15da · inbound

Validity-Calibrated Reasoning Distillation cites this paper.

Validity-Calibrated Reasoning Distillation Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T15:19:39.713044Z digest=sha256:0bafbfa5747260bbfd29eabc951e97229b1f5e27d1f0b75a7086626132e5c10c

Observation 4a4a2f6e-99c8-48e3-a917-0bfa2a78aede · inbound

Validity-Calibrated Reasoning Distillation cites this paper.

Validity-Calibrated Reasoning Distillation Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T02:31:06.308121Z digest=sha256:931955badf06762474c7b312dfa0c1e5581eb88d4c8fa0ae70d1828ef163ec79

Observation d8ca5f37-c0f4-41ab-8928-51e0d3ccc0aa · inbound

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces cites this paper.

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T02:57:15.521594Z digest=sha256:b843bef3878f6816fc1ee4e149a4aa79ecb80bd6c2a8c04671e773c38f954290

Observation 0c80e584-942d-45ed-92bb-461def9eac17 · inbound

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs cites this paper.

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T03:46:48.498800Z digest=sha256:2d7d39bf048d315664075103766ebdd1895b473e39a34ce4f43304c1e1ba5bc7

Observation 5cf0c92e-db2f-428c-8fda-99a551535990 · inbound

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs cites this paper.

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T14:31:12.591994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:31:12.591994Z digest=sha256:6b4082c8427ae5de671dfb2e03936462e18045dedfcc7b69afd8648c7797fe88

Observation d3f92752-2110-4304-a475-22c015ef5c2e · inbound

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations? cites this paper.

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations? Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:33:45.008370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:33:03.397468Z digest=sha256:514254fd2b51149c43f77287c1b6f6ec3270c570d9f062d71959be8af6f36209

Observation 1b7ea2fe-d53d-4601-a505-9f236ce6229f · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 125

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:17:24.928692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:5878bf2b818d28bb2e38a99697fdaed1e557324bf72d39f090e57294ec187482

Observation e29a27f3-24a8-4962-837b-10c619cf0093 · inbound

Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training cites this paper.

Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-03T13:18:12.265550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T13:11:44.893703Z digest=sha256:18b08f773dea5849f2a094cf2a0937a65e5154e83a58ad13356915975f09b3d6

Observation 9a1ecbb9-838e-406b-acd3-c244cf0f52fc · inbound

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information cites this paper.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:36.403464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:36.403464Z digest=sha256:2ab2a8793d0cade09682978eb03aad721399b65c0fa473b78222dd45bc108569

Observation 753a4f1e-a2c1-4cdb-ae93-c89f93ad31be · inbound

Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning cites this paper.

Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T06:14:29.370407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:14:29.370407Z digest=sha256:6945f3978f7825380b8b3da66ee9de12b08d15c0ee1418f9df3364143c56468d

Observation e5e85653-5570-4a7e-956f-0437e5891064 · inbound

EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers cites this paper.

EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:18.638135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T16:29:18.638135Z digest=sha256:5ac331326d800d5777ac2945c6af5a31a4f171b2387de5fda9c2ea7a0093aa5d

Observation feebda8d-256e-41e7-8181-d1053c14fbc5 · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.481230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.481230Z digest=sha256:2ef060972a61c361103da55360a7d053b668ed59eedb4aea0029ec16ab86ed1e