Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:08:17.955891Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 0 inbound Pith citation observations for arXiv:2506.00482.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:08:17.955891Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
83 of 83 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f5b9bc80-2b88-42ea-94a2-c81776cfe024 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2efa816f-2096-493c-9719-742628342339 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CaLMQA: Exploring culturally specific long-form question answering across 23 languages
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 427b5538-3fc3-4592-a80a-46dc79315ec1 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Program Synthesis with Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02652e39-b34d-480a-8d59-119b86945385 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Axolotl: Scalable fine-tuning framework for llms
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70cf9db4-a008-4061-b223-5319f8215fd2 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation PIQA: Reasoning about physical commonsense in natural language.Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):7432–7439, Apr
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c7944e0-cd66-499e-acff-4d0ee4f76357 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Evaluating Large Language Models Trained on Code
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e73d9e-db43-402d-8436-39b403c66e3a · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6deff245-bd09-4939-867c-5908395d103e · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 793c3c87-6d02-4ab3-aa89-120a638244cf · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2828d71c-cc47-4304-8dec-f632bc8c7838 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Training Verifiers to Solve Math Word Problems
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 072cf44b-c212-4d71-b7d3-1275b99ec56c · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 144866f4-6643-48b7-b0de-4ca0fff94548 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation SciEx: Benchmarking large language models on scientific exams with human expert grading and automatic grading
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c622398f-affd-4769-8e73-e3cb93c2f173 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c027e46-3d15-4ef3-8b6c-d084e209aa6b · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d306e540-5cc0-4a1f-b69b-a6d272c47262 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation NativQA: Multilingual Culturally-Aligned Natural Query for LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a19cbe23-d5a2-425a-beb6-719f4e207aa6 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Measuring massive multitask language understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c2d87d0-10fe-4bec-862b-7d175d674e06 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Measuring mathematical problem solving with the math dataset
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2c7b2e9-2678-4c5d-807b-1346ce605d34 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MedQA-SWE - a clinical question & answer dataset for Swedish
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91da4474-d693-434e-be15-8ab7ed1a45b5 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Liger Kernel: Efficient Triton Kernels for LLM Training
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0849f0c-eafa-4615-97f2-6f26cd597f58 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MoralBench: Moral Evaluation of LLMs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 384af186-ac2b-4525-9d81-b3fc01d6e567 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KoBBQ: Korean bias benchmark for question answering.Transactions of the Association for Computational Linguistics, 12:507–524, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5c7092a-64d8-4ce1-afb6-19c1b56f2281 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Dynabench: Rethinking benchmarking in NLP
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c1e2b82-d1b1-49cf-8110-794e6fa5f3f2 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CLIcK: A benchmark dataset of cultural and linguistic intelligence in Korean
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1db570e7-318b-486f-88f3-dddbcd66b252 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Developing a pragmatic benchmark for assessing Korean legal language understanding in large language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c49a033-ff53-4aec-b58c-66fc59956c6d · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5623cccb-be36-42ef-b4fe-07c8e3111493 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation The NarrativeQA reading comprehension challenge
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b21840f-e276-48e2-99dc-7b9e718c39b8 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KorMedMCQA: Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51d8d5d0-d31e-4d17-9557-1320c5baa8c3 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c3d54d9-ab05-4870-a72c-8534c0f1cf09 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KoSBI: A dataset for mitigating social bias risks towards safer large language model applications
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa081316-6729-4805-b54c-00c92fb35d73 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KorNAT: LLM alignment benchmark for Korean social values and common knowledge
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b5fb4e9-c1fb-4fef-982b-9a5323918bb6 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation LegalAgentBench: Evaluating LLM Agents in Legal Domain
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67f7a762-27b5-4fb1-b424-061269a1fc1d · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d952236-2223-4d8c-ad29-5afd2ade9bda · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation TruthfulQA: Measuring how models mimic human falsehoods
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9543c30d-1994-4996-b889-20b8a5679c43 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Benchmark data repositories for better benchmarking
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8efbfdfb-c22c-4d53-8960-c00821637f53 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e283af4-44c8-41ed-937e-5f530e085286 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Can a suit of armor conduct electricity? a new dataset for open book question answering
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab79f9d2-19ed-49ae-92f0-8831166b35e3 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e86d7949-42bc-49ed-886f-eb7ae1559373 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 970e71b2-80f9-4235-8a20-6bc47dc0dba2 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation BLEnD: A benchmark for llms on everyday knowledge in diverse cultures and languages
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b746cf2b-6230-403e-8f8f-4ca4330ee10b · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Extracting cultural commonsense knowledge at scale
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc5d50d0-ff3c-4ea3-9cf0-2eff798a6137 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MixEval: Deriving wisdom of the crowd from LLM benchmark mixtures
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a167dabf-4945-4a79-84f1-5217c4f6ea1f · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Chatterji, Faisal Ladhak, and Tatsunori Hashimoto
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c92b6544-3460-43fc-b53b-4942d32311cf · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation BBQ: A hand-built bias benchmark for question answering
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b115779b-e43d-467f-9b5d-3848fd76cf1a · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Survey of Cultural Awareness in Language Models: Text and Beyond
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 434e5316-32d9-4fce-937c-0ebbae8b572c · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Zero: Memory optimiza- tions toward training trillion parameter models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 431e4c29-6e41-4595-958d-42b3722aa736 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation NormAd: A framework for measuring the cultural adaptability of large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5fc97c41-d2a5-466c-9cd3-876a0506eda3 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation DiversityMedQA: A benchmark for assessing demographic biases in medical diagnosis using large language models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a931e440-3305-4057-84b3-49e2394adb4d · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29c42e3f-8d2a-4402-9d66-c368ccbf24cb · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Kochenderfer
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32256750-95a7-49dc-bd84-fbcc2c0f5f5e · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation WinoGrande: an adversarial winograd schema challenge at scale.Commun
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc6416ba-d929-4711-95aa-cd19b259b7d7 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Social IQa: Commonsense reasoning about social interactions
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 597877f1-8911-4dba-b89d-f29fbcd96adb · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Benchmarks as microscopes: A call for model metrology
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7db6d897-08bd-47d6-854d-be04c1ce887e · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Multi-fact: Assessing factuality of multilingual llms using factscore, 2024
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d3478be-d6f6-4cab-b9db-7783e1a7ad1d · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation YourBench: Easy Custom Evaluation Sets for Everyone
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac95e546-01e9-4a9e-b94c-0804e1c866ba · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CultureBank: An online community-driven knowledge base towards culturally aware language technologies
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2632e0cc-e035-437f-9565-3dfdb648bb34 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26e37416-51ce-415c-9642-e621f807dab6 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KRX bench: Automating financial benchmark creation via large language models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad0d7d1a-0d36-4340-9651-9a858d447735 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Beyond Classification: Financial Reasoning in State-of-the-Art Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3201edf1-4f3e-4d99-a74a-b41de35e9f1c · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Multi-step reasoning in Korean and the emergent mirage
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 624cdaba-8ce7-4ba1-b8f4-cfe3cd287df1 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KMMLU: Measuring massive multitask language understanding in Korean
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 062acff7-3666-48da-83a3-d3376048856b · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation HAE-RAE bench: Evaluation of Korean knowledge in language models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51cdb724-54f3-47c5-9fc7-280e106df693 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Challenging BIG-bench tasks and whether chain-of-thought can solve them
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7a0cb8e-86be-4b03-ac03-4c9eb57a57ad · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 432bf232-706d-4036-b311-207966c0a5a0 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CommonsenseQA: A question answering challenge targeting commonsense knowledge
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00d892b9-90c2-4f7b-a54b-7846b57cd4c5 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Gemma 3 Technical Report
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b644f29-8f98-4946-a501-c1d937aec4bd · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Benchmark suites instead of leaderboards for evaluating AI fairness.Patterns, 5(11):101080, 2024
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6098c32e-3b67-48eb-92a2-d20244f6f03a · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation SeaEval for multilingual foundation models: From cross-lingual alignment to cultural reasoning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1fd0d21e-54f8-4e60-83f6-9b48a90b3a76 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49db165a-1c43-4554-97df-544abd4e6381 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af13bd18-0319-4209-a00e-594f1aa283f0 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MMLU-Pro: A more robust and challenging multi-task language understanding benchmark
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f44943f3-cb76-4bfa-aeeb-bf86bf08a589 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Toward an Evaluation Science for Generative AI Systems
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3199245-1f8c-4d78-958b-bed6e0247619 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Qwen3 Technical Report
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2d4076c-8373-42b7-bbb3-a99321d88f7a · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Qwen2.5 Technical Report
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a0588e0-0c37-4bd2-95fb-f1c7e168d36e · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc172362-c9b7-4ef7-9b97-e23956a44a2b · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation FLASK: Fine-grained language model eval- uation based on alignment skill sets
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7da61289-4ca5-42d3-86b8-83542a2efdc6 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation GeoM- LAMA: Geo-diverse commonsense probing on multilingual pre-trained language models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 725c2996-fa56-420c-b4f4-d17e36ae7af6 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad13bf2f-b95b-429a-96e4-7f7b91ef36e0 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Maruf Hossain, Guang- Jie Ren, Kate Soule, Yifan Mai, and Yada Zhu
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29a84389-3535-41e3-85c8-51ccf4318ed9 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Task me anything
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86d268ff-e827-4c90-a2e4-463393160202 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Users can interactively explore the overall data distribution they are interested in
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e0de9c67-b6cb-47e7-acf9-528c7ec99ed7 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation By reviewing samples, users can verify whether the dataset matches their needs and explore datasets suitable for their purposes
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f2adf41-922e-4c81-9f4c-6dc5c22c83cd · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Is the Earth flat?
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88c7fafc-7f12-49be-82bc-bc53c5530e79 · outbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.