Pith. sign in

Paper Citation Record · LEDGER

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

As of 6 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 33 inbound Pith citation observations for arXiv:2305.12474.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.12474 v3

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T12:28:32.395213Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:21:25.985017Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

19
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5da42aed-6caf-46c4-a61f-bfa49b6d6a31 · outbound

This paper cites Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.495884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:6cd88a818587ff492e9a42ff4487167353b069e11cc1dc262386fc70554ccdc8

Observation 1a8e6288-784b-4c65-83d3-4f31f40cdeb7 · outbound

This paper cites GPT-4 Technical Report.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark GPT-4 Technical Report

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T12:28:32.420686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:569b899d1b5d234562b65b43a4b0e56ed354d0cb6c3c87d0de23625d48993131

Observation c48315cd-a1d6-454d-9eb1-1eb430b1ae93 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T12:28:32.430613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:2584f8b1f457c6191a50f48f0665439353837b48a3202ea059d9235906267562

Observation 2e12e333-9264-4843-8436-8c52a50f0f68 · outbound

This paper cites an unresolved cited work.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-17T12:28:32.507446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:12096a357fa7b1e29af7c2a1a8db98f37f1fad32c04c3236d12fc36106b98b83

Observation 8ee23b33-4950-482b-9801-a2a255f74204 · outbound

This paper cites an unresolved cited work.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-17T12:28:32.437185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:ee34543080b14f6b9ca2fa6d507c22c1258e6c0547571f6240b30e545a0eace6

Observation 7f894159-50dc-4532-88d9-738449ee9e38 · outbound

This paper cites an unresolved cited work.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-17T12:28:32.441022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:f63d8e836de6d83bd65f4a8cd0a2dd1a4097ca7a30e87cf5f7207aa32aff4751

Observation 2dc7f7f6-d212-4254-919e-b62a1eb59f3d · outbound

This paper cites In order to protect this heritage while also developing tourism activities, measures need to be taken to protect the tourism resources.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark In order to protect this heritage while also developing tourism activities, measures need to be taken to protect the tourism resources

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.446123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:824de90c07cf228454956da45312def31c73b152529756fc39d8aab20a64f9b2

Observation fdace3f6-08ac-41dd-ad0a-e5043aa0912e · outbound

This paper cites an unresolved cited work.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-17T12:28:32.449944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:9f3b7f4e5daae8794e6ae4935e581c07903472ee5618f50d89b380339508a423

Observation 4e33bdc9-1f4b-4351-98dc-0e868daf10c2 · outbound

This paper cites This will enhance the cultural literacy and environmental awareness of the tourists and reduce the damage to the terraces.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark This will enhance the cultural literacy and environmental awareness of the tourists and reduce the damage to the terraces

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.453796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:4dc5852a946b76166e865bd3f4a0a071a58b20bc05a8cbbe06a1873ee1b49475

Observation 39dada78-c3c9-4183-869c-eca92b621828 · outbound

This paper cites an unresolved cited work.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-17T12:28:32.457230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:a8b5f0b4b39cb2870139ad14306792912b91123bd9fecf46e0c46bf65d5cc593

Observation 7cd947d7-4ba4-49b6-88d3-2d4f26dc2c62 · outbound

This paper cites At the same time, these facilities should be planned judiciously to avoid damage to the terraces.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark At the same time, these facilities should be planned judiciously to avoid damage to the terraces

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.462517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:447e39cfcb406296a6078dda28517882b3918fead7bfbac88c1f45720f2357e5

Observation a3ac400a-ca29-4c75-9286-86e27fc02386 · outbound

This paper cites 完善 景区规划、依法保护生态环境.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark 完善 景区规划、依法保护生态环境

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.466212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:fd55ca69e6d7197df843aec7386d74f68834e7059bc019b069e451fdd8fb858a

Observation d84ba65c-aabf-45ed-ae82-c06c3f386dff · outbound

This paper cites 普及旅游文化环境保 护教育,提高游客对旅游资源环境保护 的意识.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark 普及旅游文化环境保 护教育,提高游客对旅游资源环境保护 的意识

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.469890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:eaebff3cbe7f3dafdbdbab738803f3d5b4758ef3ad3408515cbc3cb847b5a557

Observation 299c4ed3-e7e0-43e1-a9a8-f4b8798faa69 · outbound

This paper cites 评定该‘生 态博物馆’的环境容量,对人口数量的 容纳程度,限制客流量.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark 评定该‘生 态博物馆’的环境容量,对人口数量的 容纳程度,限制客流量

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.473506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:79660a53573cf2e3309371a993ea783f6eac872d3d190922b54bd75d1267ae90

Observation b92bfb5d-8eac-4c8e-86f8-4e7640a18f7c · outbound

This paper cites 尽可能保证新建设施与景区景观相 融合.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark 尽可能保证新建设施与景区景观相 融合

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.478599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:2aa30511f805bd40cd932eb2e5f81026e03d7d311dae76af047fc12fbc10d611

Observation 113cf6b0-eb19-4257-a3b1-60c8a18062e8 · outbound

This paper cites improve the planning of the scenic area, protect the ecological environment in accordance with the law.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark improve the planning of the scenic area, protect the ecological environment in accordance with the law

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.482604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:d5a3fb0c367085f232b57e05f0179a4fc6a4a9c6266aeebaf8f07e30c64766a6

Observation 2c8949de-7100-4f53-a883-3c03366751a6 · outbound

This paper cites popularize education on the protection of the tourism cultural environment, raise tourists’ awareness of the protection of tourism resources and environment.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark popularize education on the protection of the tourism cultural environment, raise tourists’ awareness of the protection of tourism resources and environment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.487677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:8568c557a270b25017a203f87fe7ad10ea833b9d445defeb00a5dd590a2696ca

Observation 920ba379-6e08-437c-acc4-4c22993a1574 · outbound

This paper cites assess the environmental capacity of this ‘Ecological Museum’, regulate the carrying capacity in terms of population, limit the flow of visitors.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark assess the environmental capacity of this ‘Ecological Museum’, regulate the carrying capacity in terms of population, limit the flow of visitors

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.491874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:13213bcab8a985c6256f18022805954f0d0790903772e16f56159c7bbb317fe0

Observation ea8ba8d1-5dab-498d-99f4-bc6d9ff60ae4 · outbound

This paper cites ensure new facilities blend harmoniously with the scenic landscape.

Evaluating the Performance of Large Language Models on GAOKAO Benchmark ensure new facilities blend harmoniously with the scenic landscape

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:28:32.503672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T12:28:32.395213Z digest=sha256:5ff043571a73a9d53ea3150c3ef816dfc32bd1ebc3981dc434342d656c3d5772

Pith citing papers

Observation 42c25863-cfe0-47a6-94c5-2b7f29e04c1c · inbound

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism cites this paper.

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T06:08:05.550346Z digest=sha256:21248a18457a013b585fba3b3272c09ce81a9172c7fc3700c8d90b5bdff3973a

Observation 1cd82ad0-a844-4020-973b-484bf5de6693 · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:f863ed7441996a0d2fb2a1d3c213b68f2ae9ce746c86b232cfe9f244cfea6f58

Observation 0ec704b1-b80c-474d-926c-be559c777325 · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:fbcc9a2005ccb41e3110f6db14ccb141546269695fd679a50168bccd3e9d8952

Observation ae6bd359-f7f9-4914-8730-115862034a40 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:b6efc3d03e54749a0fabff9dc83d88905b5249edbc4ade498a140147241d5300

Observation 8613c96b-2375-4646-93c4-24db5a62c5c3 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:2e6ff4eba27984cff6d3e79bf683dacf162915eec07f910f12975b37f167a6a6

Observation 3f5a99a0-ade0-42fc-a449-04d9131a51f9 · inbound

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource cites this paper.

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:05:47.729661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:05:08.916339Z digest=sha256:80dc85953a9bcf3f0580d1770b6799e85b51ff514f494aaa885d5a4f40db93a8

Observation 03dd5117-4e25-478c-8bb9-bfd808dda126 · inbound

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning cites this paper.

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 207

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:01:10.027080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T01:01:09.840919Z digest=sha256:d1e2dcedc3d1cf000c3413f4aeec9b7487ac46a70881867176001da1f7e6af71

Observation e9f3f8a1-0dd0-48bb-80cc-36c9b19869cf · inbound

TASE: Token Awareness and Structured Evaluation for Multilingual Language Models cites this paper.

TASE: Token Awareness and Structured Evaluation for Multilingual Language Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T23:21:25.985017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:21:25.985017Z digest=sha256:5c7e59c8dcefec72c8edec16627318381ac2329ac6c3ae6ad8da74a1c6592b8a

Observation 310061f6-30b1-4ac3-ba9b-a8151a1c8305 · inbound

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support cites this paper.

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-05T23:04:57.082913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:04:57.082913Z digest=sha256:5734e77fdd05383364fde675fde542024f6a9e9653ec2546fcb1e4ff0ae9b4b6

Observation dac95634-5e48-4fae-897d-b95e50e70bcc · inbound

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models cites this paper.

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T18:53:05.214761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:53:05.214761Z digest=sha256:17da51d0a29364fc5d8f0993229ed7538a27354a3f1ae3c61e81d77f6d5847b7

Observation ca6715dc-d740-4dac-ab4f-437a436c501e · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:301c11f8445db0a3b12b02c399402cec0804e9280f15d9c80986a0c54d2719b0

Observation 6ac50c29-e589-40b4-bb08-54936b5c90b9 · inbound

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards cites this paper.

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:48.547800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:15:48.547800Z digest=sha256:15e920a59a2a26f6fa268bbace6091dd72e179baf5cf9705d147f3ac9e5659fa

Observation ea268b62-2206-4296-8766-bdf9dc750395 · inbound

LLaDA2.0: Scaling Up Diffusion Language Models to 100B cites this paper.

LLaDA2.0: Scaling Up Diffusion Language Models to 100B Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:53:20.911374Z digest=sha256:14ae089fcc4770ed114229946c4f1039ef2ed2effad6c5363fd89518c7804145

Observation b598bf54-1b21-4e92-a36c-aca177dcc8e8 · inbound

GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents cites this paper.

GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:05:26.681888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:02:28.660002Z digest=sha256:13bc0224c1cb910845963aa5ef56b18c41173ac5105eeb640bdc29626aa6a012

Observation 55e59d7a-4086-49e2-9e72-39ecb8321610 · inbound

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding cites this paper.

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T09:11:31.870441Z digest=sha256:61ec623eb42ed65ae635be16c05f613e3798bfb6236ebed7f7bc21b943320e06

Observation 35fe28d5-b4bb-45b2-a6f0-ccbf62ff98e3 · inbound

TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice cites this paper.

TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T17:59:44.844149Z digest=sha256:c968a74c2d62405cdefa365cd12395390c609eb24fee350280c2064b854f259b

Observation 92b3a011-13f8-41b1-980d-0c82dc16e045 · inbound

SFT-GRPO Data Overlap as a Post-Training Hyperparameter for Autoformalization cites this paper.

SFT-GRPO Data Overlap as a Post-Training Hyperparameter for Autoformalization Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T13:21:52.225115Z digest=sha256:22099832d428c1adb7796deda41b9dc1ab4275c7212a91cbcef9440bea46e10e

Observation eabcf859-5c2b-4db3-8fa3-4864e67a1e31 · inbound

RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025) cites this paper.

RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025) Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 33

Resolution
malformed identifier
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T22:01:21.565986Z digest=sha256:9d38d4112eb924466c3f843d9d4bb15eada3d64cd2e148cf175f0ac7f9b46777

Observation 5cdbe6e8-0a71-4fde-86f8-073a2554e6d1 · inbound

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs cites this paper.

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T00:51:37.506096Z digest=sha256:a561e20a114fbd7324af972d08244096b3b4c714aef3ede228ecf21e38123575

Observation 534afaf3-715d-452a-b419-7f00bad02d07 · inbound

When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling cites this paper.

When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T11:00:21.413246Z digest=sha256:cc2ddc4158864001f739d58670a326efc93cfe4cab6fae778796291d18dd341f

Observation 456f76f0-19c1-4c20-8436-2b23c688c80b · inbound

When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling cites this paper.

When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T05:23:31.540079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:23:31.540079Z digest=sha256:3a1d638213ad92870fbabb328f6984b628a18162697905eb56c3858cc8fb6c92

Observation 70da74c3-a614-47bb-909c-1164359e15da · inbound

Validity-Calibrated Reasoning Distillation cites this paper.

Validity-Calibrated Reasoning Distillation Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T15:19:39.713044Z digest=sha256:e5151c0cb5c89e15921b8cdf3a855c778e25bcd0edd056f345cab4ca59173f70

Observation 4a4a2f6e-99c8-48e3-a917-0bfa2a78aede · inbound

Validity-Calibrated Reasoning Distillation cites this paper.

Validity-Calibrated Reasoning Distillation Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:31:06.308121Z digest=sha256:db684e5efcd7e7868ca7ff74e7644c3874a945f2848aba1eb104c29f8b62e39d

Observation d8ca5f37-c0f4-41ab-8928-51e0d3ccc0aa · inbound

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces cites this paper.

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T02:57:15.521594Z digest=sha256:12ee527c8eb0fb1614ae4583ab6dfd3bd5529be86e11920ade5f4ef1e05a1e13

Observation 0c80e584-942d-45ed-92bb-461def9eac17 · inbound

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs cites this paper.

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:46:48.498800Z digest=sha256:8a57eba8aa8c9982560910bf22fda1094d2f6c2b376c9618ff22959110f33c3e

Observation 5cf0c92e-db2f-428c-8fda-99a551535990 · inbound

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs cites this paper.

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T14:31:12.591994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:31:12.591994Z digest=sha256:800271841b8f7a1056dca225a5333f640c9a182e87165ace78e5b7adff7042d5

Observation d3f92752-2110-4304-a475-22c015ef5c2e · inbound

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations? cites this paper.

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations? Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:33:45.008370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T17:33:03.397468Z digest=sha256:b190921a0c6869d058d6fe873165b2734653facf1c54b623d46375e362b324c3

Observation 1b7ea2fe-d53d-4601-a505-9f236ce6229f · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 125

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:17:24.928692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:7efdb7fa064caf8d46e5fe9d3fe9a17a3a06ac4aa6e7ebb463d67158cd9cbeae

Observation e29a27f3-24a8-4962-837b-10c619cf0093 · inbound

Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training cites this paper.

Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-03T13:18:12.265550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T13:11:44.893703Z digest=sha256:d56c3a903bf90ded3a2da9fa077e9a221d95ddfd284e622556e87f4b3745e370

Observation 9a1ecbb9-838e-406b-acd3-c244cf0f52fc · inbound

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information cites this paper.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:36.403464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:36.403464Z digest=sha256:023bc5b47dd60526316a18304a265337b6def783569db805e25d834ec6398299

Observation 753a4f1e-a2c1-4cdb-ae93-c89f93ad31be · inbound

Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning cites this paper.

Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T06:14:29.370407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:14:29.370407Z digest=sha256:1731a7408584b3a65b590a9c540bed187e5d91b72d440e74dad3807660d33591

Observation e5e85653-5570-4a7e-956f-0437e5891064 · inbound

EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers cites this paper.

EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:18.638135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T16:29:18.638135Z digest=sha256:e63e5885b61a209b51bfb97ae09929381ab7a14134838e5a6e49bc0aae801e46

Observation feebda8d-256e-41e7-8181-d1053c14fbc5 · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.481230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.481230Z digest=sha256:8e2639efa0d6b92e0cfc8b35806748ede9e80bdc6051cdbbfdbbcb9e9abc7120