Pith. sign in

Paper Citation Record · LEDGER

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation

As of 7 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 0 inbound Pith citation observations for arXiv:2506.00482.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00482 v1

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:08:17.955891Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

83 of 83 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5b9bc80-2b88-42ea-94a2-c81776cfe024 · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:09.977304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:09.977304Z digest=sha256:7f37197adcd46c3dcfb52bd13584545a4afc73c73b110c8dc45fc0b1500058d5

Observation 2efa816f-2096-493c-9719-742628342339 · outbound

This paper cites CaLMQA: Exploring culturally specific long-form question answering across 23 languages.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CaLMQA: Exploring culturally specific long-form question answering across 23 languages

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.073701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.073701Z digest=sha256:339a1942feac846252a48ba12a02f8a4fdd4fc661dca651e8f9adb6f16f55f24

Observation 427b5538-3fc3-4592-a80a-46dc79315ec1 · outbound

This paper cites Program Synthesis with Large Language Models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.157068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.157068Z digest=sha256:513b3c9988cd0ab2a9cfa8c6bbc73d358d2c742ae8e93efa5dccffd7114d49dd

Observation 02652e39-b34d-480a-8d59-119b86945385 · outbound

This paper cites Axolotl: Scalable fine-tuning framework for llms.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Axolotl: Scalable fine-tuning framework for llms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.225114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.225114Z digest=sha256:0e463ea53998cd80ea5dd1c06fd328d9319b6f65fd02dda3e05d01546d09ff40

Observation 70cf9db4-a008-4061-b223-5319f8215fd2 · outbound

This paper cites PIQA: Reasoning about physical commonsense in natural language.Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):7432–7439, Apr.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation PIQA: Reasoning about physical commonsense in natural language.Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):7432–7439, Apr

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.311875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.311875Z digest=sha256:bf3a98b6c48ecbd5adc566782d0f58fe09168cc6eaea82a8359d92855b6724fb

Observation 7c7944e0-cd66-499e-acff-4d0ee4f76357 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.404642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.404642Z digest=sha256:1b8fceb726b9cc5cdcd9c3bd57bdd38e277dd63cf90238dce65beecf5b028a0c

Observation f7e73d9e-db43-402d-8436-39b403c66e3a · outbound

This paper cites Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:28.925972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:10.504819Z digest=sha256:fc05fac694a068569e8790cc8972b30bcdd4548d05d12796f4f198d0395ced61

Observation 6deff245-bd09-4939-867c-5908395d103e · outbound

This paper cites CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.604298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.604298Z digest=sha256:8eb8d486223c9e4a7467c9b70bd52d559c9a0a7080581acdeb2402ba31964b6c

Observation 793c3c87-6d02-4ab3-aa89-120a638244cf · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.688904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.688904Z digest=sha256:a98bd1401b04238856ca22df89c3be5a91cc5f8fe747c6e0d464fddc87a3c129

Observation 2828d71c-cc47-4304-8dec-f632bc8c7838 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.759064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.759064Z digest=sha256:c877922f4f4f063a7ec50d38d8edda253fb518e80a24fbd9c1792759f80645b8

Observation 072cf44b-c212-4d71-b7d3-1275b99ec56c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.838533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.838533Z digest=sha256:3caeaf795f91727dd5530bcfbe6cf763af6f1f597e9167fe92903d0820c56925

Observation 144866f4-6643-48b7-b0de-4ca0fff94548 · outbound

This paper cites SciEx: Benchmarking large language models on scientific exams with human expert grading and automatic grading.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation SciEx: Benchmarking large language models on scientific exams with human expert grading and automatic grading

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:28.653087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:10.913333Z digest=sha256:92a08165b23582536768d369ccb65b9803f25eb23b20ed1e4bd6ea9cdf2475a3

Observation c622398f-affd-4769-8e73-e3cb93c2f173 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.040928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.040928Z digest=sha256:c33c8990783c150871a83bfcf7c3958e40aaa7fb34a8c50d63f234c75af3e222

Observation 8c027e46-3d15-4ef3-8b6c-d084e209aa6b · outbound

This paper cites The Llama 3 Herd of Models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.143287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.143287Z digest=sha256:32e95a2dd5110360902dfe0565800e682d69723369af944417ae6e1ad042ac48

Observation d306e540-5cc0-4a1f-b69b-a6d272c47262 · outbound

This paper cites NativQA: Multilingual Culturally-Aligned Natural Query for LLMs.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation NativQA: Multilingual Culturally-Aligned Natural Query for LLMs

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:08:18.380004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:11.225092Z digest=sha256:ccc3080eda46b34db7d139238cc7eaf52af4ad2ab7efb9f1ee68e0139a9a563a

Observation a19cbe23-d5a2-425a-beb6-719f4e207aa6 · outbound

This paper cites Measuring massive multitask language understanding.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Measuring massive multitask language understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.314107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.314107Z digest=sha256:ae7b449d6b834bef16a1ce26b4a1e92e11fc197310c683827f08c74940a3e5b6

Observation 5c2d87d0-10fe-4bec-862b-7d175d674e06 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Measuring mathematical problem solving with the math dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.398525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.398525Z digest=sha256:1509e691e68b19da63511e848632de125a6fba09fb2516885559fadf45568776

Observation e2c7b2e9-2678-4c5d-807b-1346ce605d34 · outbound

This paper cites MedQA-SWE - a clinical question & answer dataset for Swedish.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MedQA-SWE - a clinical question & answer dataset for Swedish

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:28.428624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:11.541072Z digest=sha256:d4f5101f975dec7a7a5d06cf090e779c55a454d231a36a0c69c27dc337d9cfa1

Observation 91da4474-d693-434e-be15-8ab7ed1a45b5 · outbound

This paper cites Liger Kernel: Efficient Triton Kernels for LLM Training.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Liger Kernel: Efficient Triton Kernels for LLM Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.651962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.651962Z digest=sha256:34050f36778045d40a394762e8277267dca96b73d43affea1e376571f60d3bad

Observation b0849f0c-eafa-4615-97f2-6f26cd597f58 · outbound

This paper cites MoralBench: Moral Evaluation of LLMs.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MoralBench: Moral Evaluation of LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.740741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.740741Z digest=sha256:7d001662838d2dc06a41011b4fd531929ef5cf820b1994f888270506bb424c1a

Observation 384af186-ac2b-4525-9d81-b3fc01d6e567 · outbound

This paper cites KoBBQ: Korean bias benchmark for question answering.Transactions of the Association for Computational Linguistics, 12:507–524, 2024.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KoBBQ: Korean bias benchmark for question answering.Transactions of the Association for Computational Linguistics, 12:507–524, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:28.176490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:11.856041Z digest=sha256:3351c87eb4c2e78fa996d1fe6299a8864c86aaf5cc48c9b921349c7667b45324

Observation b5c7092a-64d8-4ce1-afb6-19c1b56f2281 · outbound

This paper cites Dynabench: Rethinking benchmarking in NLP.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Dynabench: Rethinking benchmarking in NLP

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:27.926812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:11.960334Z digest=sha256:9ee817a707fd94ecc1929b2dd101022cc7b958223ee73c323951b177d1f4338c

Observation 7c1e2b82-d1b1-49cf-8110-794e6fa5f3f2 · outbound

This paper cites CLIcK: A benchmark dataset of cultural and linguistic intelligence in Korean.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CLIcK: A benchmark dataset of cultural and linguistic intelligence in Korean

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:27.623537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:12.098448Z digest=sha256:705969a96e38a5a7cddb3fed979f01a659c2a65c110862808795dd770e6da70e

Observation 1db570e7-318b-486f-88f3-dddbcd66b252 · outbound

This paper cites Developing a pragmatic benchmark for assessing Korean legal language understanding in large language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Developing a pragmatic benchmark for assessing Korean legal language understanding in large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:27.371801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:12.199734Z digest=sha256:a19ca2cdaca380110f9ac345ed4b67aea5b3676d817660fcd45a99615d56c03c

Observation 8c49a033-ff53-4aec-b58c-66fc59956c6d · outbound

This paper cites Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:12.304393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:12.304393Z digest=sha256:89dba3e971b2f8f9f83460a4f7a9be7f022a15095dcb5db0f0a503e4ac2f616c

Observation 5623cccb-be36-42ef-b4fe-07c8e3111493 · outbound

This paper cites The NarrativeQA reading comprehension challenge.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation The NarrativeQA reading comprehension challenge

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:27.027075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:12.414934Z digest=sha256:ec6da6d04b97d0a087dc6696c8313339f7a4330ee52478d83a41b129185ff768

Observation 7b21840f-e276-48e2-99dc-7b9e718c39b8 · outbound

This paper cites KorMedMCQA: Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KorMedMCQA: Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:12.530994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:12.530994Z digest=sha256:6cbded4fb00942cfcac941ca9a0fb97f16708e8f8f21796ca628a05e49797c6a

Observation 51d8d5d0-d31e-4d17-9557-1320c5baa8c3 · outbound

This paper cites Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:12.613938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:12.613938Z digest=sha256:f234b20d6e1c3c97ac347266679ed8bc5ba354066f176362b5ed37c7476c0119

Observation 6c3d54d9-ab05-4870-a72c-8534c0f1cf09 · outbound

This paper cites KoSBI: A dataset for mitigating social bias risks towards safer large language model applications.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KoSBI: A dataset for mitigating social bias risks towards safer large language model applications

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:26.686665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:12.725096Z digest=sha256:f2e707cf8a9c8892472036fbb2cf4451a6b51784839b994694e844ed0afa68b6

Observation aa081316-6729-4805-b54c-00c92fb35d73 · outbound

This paper cites KorNAT: LLM alignment benchmark for Korean social values and common knowledge.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KorNAT: LLM alignment benchmark for Korean social values and common knowledge

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:26.394750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:12.848365Z digest=sha256:f8aebd55070166ce17c7ee206bfe32cf7b647df67b6610df3a737721849d2576

Observation 6b5fb4e9-c1fb-4fef-982b-9a5323918bb6 · outbound

This paper cites LegalAgentBench: Evaluating LLM Agents in Legal Domain.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation LegalAgentBench: Evaluating LLM Agents in Legal Domain

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:12.933539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:12.933539Z digest=sha256:92804e0d6e1aaf0968c8834219b0cffc61b54bf8e2c0f975266e54396d3fec1b

Observation 67f7a762-27b5-4fb1-b424-061269a1fc1d · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:08:26.087765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:13.031385Z digest=sha256:49e754b157d974cbb52edbc09828865c950b6ba0b3fd55fb4cae0a878f559860

Observation 0d952236-2223-4d8c-ad29-5afd2ade9bda · outbound

This paper cites TruthfulQA: Measuring how models mimic human falsehoods.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation TruthfulQA: Measuring how models mimic human falsehoods

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:13.117735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:13.117735Z digest=sha256:e1b21cf97e13914947f0be3d1ebc1a997ddf1db9da1e75de2d178d47b626e5cd

Observation 9543c30d-1994-4996-b889-20b8a5679c43 · outbound

This paper cites Benchmark data repositories for better benchmarking.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Benchmark data repositories for better benchmarking

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:25.844219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:13.193135Z digest=sha256:47b851a6001d482e0a517c2dffe4f64986d21f5727793410bb7f75b99d03c58c

Observation 8efbfdfb-c22c-4d53-8960-c00821637f53 · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:08:25.663250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:13.299417Z digest=sha256:857f9fb4001830262cb2dd334b6b1a8fc60ee796cd8009920f38bf553b851c23

Observation 3e283af4-44c8-41ed-937e-5f530e085286 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:25.342672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:13.416530Z digest=sha256:05a1013d247c28ab139564a21f3787477d8cf307bd703a8f7c8247f2ac922834

Observation ab79f9d2-19ed-49ae-92f0-8831166b35e3 · outbound

This paper cites FActScore: Fine-grained atomic evaluation of factual precision in long form text generation.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation FActScore: Fine-grained atomic evaluation of factual precision in long form text generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:25.078670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:13.497724Z digest=sha256:acc2e900131d1272846dcd838488547d5295ea0d06c985951428b3b3c8d69bbd

Observation e86d7949-42bc-49ed-886f-eb7ae1559373 · outbound

This paper cites Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:13.563546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:13.563546Z digest=sha256:1bc2ed8c4a9b908befe736983e0db06c3e1b99458777c655b8e4f81fe85e806b

Observation 970e71b2-80f9-4235-8a20-6bc47dc0dba2 · outbound

This paper cites BLEnD: A benchmark for llms on everyday knowledge in diverse cultures and languages.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation BLEnD: A benchmark for llms on everyday knowledge in diverse cultures and languages

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.796108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:13.674542Z digest=sha256:8d6881797035e7ddb3bf5c48b48423fbab71e927a3feb19a2abd90a5b2d4ed91

Observation b746cf2b-6230-403e-8f8f-4ca4330ee10b · outbound

This paper cites Extracting cultural commonsense knowledge at scale.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Extracting cultural commonsense knowledge at scale

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.597620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:13.809086Z digest=sha256:8a4e7373d37f94489599027631a4b0f89c8381bf6b03587b539693941d26af5a

Observation dc5d50d0-ff3c-4ea3-9cf0-2eff798a6137 · outbound

This paper cites MixEval: Deriving wisdom of the crowd from LLM benchmark mixtures.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MixEval: Deriving wisdom of the crowd from LLM benchmark mixtures

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.233222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:13.880832Z digest=sha256:8c8c68c60538df6833756544637266e0d7cef84446768062fa51193911419d24

Observation a167dabf-4945-4a79-84f1-5217c4f6ea1f · outbound

This paper cites Chatterji, Faisal Ladhak, and Tatsunori Hashimoto.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Chatterji, Faisal Ladhak, and Tatsunori Hashimoto

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:13.986625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:13.986625Z digest=sha256:71c70b5551f04271d512b16437ec6afec8b34c1fce78e4d7cad8e91e59927d66

Observation c92b6544-3460-43fc-b53b-4942d32311cf · outbound

This paper cites BBQ: A hand-built bias benchmark for question answering.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation BBQ: A hand-built bias benchmark for question answering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.031957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:14.078274Z digest=sha256:3a9093e866918cba6ebecff24c8e79f506904ae156d4374791b64e8e366f99fa

Observation b115779b-e43d-467f-9b5d-3848fd76cf1a · outbound

This paper cites Survey of Cultural Awareness in Language Models: Text and Beyond.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Survey of Cultural Awareness in Language Models: Text and Beyond

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:14.144191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:14.144191Z digest=sha256:a9bd1c8e5f425d43e7c90f56f95ae1f60f8031fd0b61c9315f435d7c3e5f47db

Observation 434e5316-32d9-4fce-937c-0ebbae8b572c · outbound

This paper cites Zero: Memory optimiza- tions toward training trillion parameter models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Zero: Memory optimiza- tions toward training trillion parameter models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.722060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:14.231714Z digest=sha256:ca347539bedd788ade13ab09c83b66520047435efc8aaaf25a2026945fd642b6

Observation 431e4c29-6e41-4595-958d-42b3722aa736 · outbound

This paper cites NormAd: A framework for measuring the cultural adaptability of large language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation NormAd: A framework for measuring the cultural adaptability of large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.401621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:14.335556Z digest=sha256:78bce3d74171b829364e222eb46154bbecf532b3067254c69b8f96c279411aeb

Observation 5fc97c41-d2a5-466c-9cd3-876a0506eda3 · outbound

This paper cites DiversityMedQA: A benchmark for assessing demographic biases in medical diagnosis using large language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation DiversityMedQA: A benchmark for assessing demographic biases in medical diagnosis using large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.085948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:14.581006Z digest=sha256:276f2d7dcffd0b87f5b44ddcee6f4ec073b4370c8c5dd3d861f62388c0cb2173

Observation a931e440-3305-4057-84b3-49e2394adb4d · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:08:22.795446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:14.691459Z digest=sha256:c8f3e178ea13ae78096466c9b12f1efac5bc35a543bd603a7d9e16c116862fa3

Observation 29c42e3f-8d2a-4402-9d66-c368ccbf24cb · outbound

This paper cites Kochenderfer.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Kochenderfer

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.550842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:14.833051Z digest=sha256:710ac48f5949f365ceb5aa359b166c61ced28673e69f4f85aef191fc66a6b387

Observation 32256750-95a7-49dc-bd84-fbcc2c0f5f5e · outbound

This paper cites WinoGrande: an adversarial winograd schema challenge at scale.Commun.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation WinoGrande: an adversarial winograd schema challenge at scale.Commun

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.369793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:14.924761Z digest=sha256:977c59ab4d6e3d3add0661b6f842cda743440fce766a5c76d8c7b64baf9f8b5b

Observation cc6416ba-d929-4711-95aa-cd19b259b7d7 · outbound

This paper cites Social IQa: Commonsense reasoning about social interactions.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Social IQa: Commonsense reasoning about social interactions

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.146589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:15.015768Z digest=sha256:b023892a92b62afd97eb394a36672d8da24dc1ec98cab929274747127d1ed3a0

Observation 597877f1-8911-4dba-b89d-f29fbcd96adb · outbound

This paper cites Benchmarks as microscopes: A call for model metrology.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Benchmarks as microscopes: A call for model metrology

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.940061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:15.107192Z digest=sha256:255af1fefa4c6dff3f765fa4dea8569109f765644de419f437387df3b659d931

Observation 7db6d897-08bd-47d6-854d-be04c1ce887e · outbound

This paper cites Multi-fact: Assessing factuality of multilingual llms using factscore, 2024.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Multi-fact: Assessing factuality of multilingual llms using factscore, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.694305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:15.229945Z digest=sha256:358901bcbfef9dd19ecdc2bc214526ce7b3e27d226094988e42a07c41c06da7a

Observation 1d3478be-d6f6-4cab-b9db-7783e1a7ad1d · outbound

This paper cites YourBench: Easy Custom Evaluation Sets for Everyone.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation YourBench: Easy Custom Evaluation Sets for Everyone

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:15.321570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:15.321570Z digest=sha256:7fa0b97a4ee2c87c3b3d1f38f43da56e6ae7546857644d0421945fa7f2dc92d8

Observation ac95e546-01e9-4a9e-b94c-0804e1c866ba · outbound

This paper cites CultureBank: An online community-driven knowledge base towards culturally aware language technologies.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CultureBank: An online community-driven knowledge base towards culturally aware language technologies

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.384899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:15.440338Z digest=sha256:7baa7fa7346869acbdf265a8d2e10afa665d9ce90ccb2ab44aeb0ad7dceac158

Observation 2632e0cc-e035-437f-9565-3dfdb648bb34 · outbound

This paper cites Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:15.542569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:15.542569Z digest=sha256:7381fd2152986dbc7a9c11f3901c31a97ed62f945eb77a181c16b246642ae84f

Observation 26e37416-51ce-415c-9642-e621f807dab6 · outbound

This paper cites KRX bench: Automating financial benchmark creation via large language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KRX bench: Automating financial benchmark creation via large language models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.170385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:15.639270Z digest=sha256:3f1ddf50e646a2752ce321179747c1f2d5b23e523042b04fdaeaf0df06e43b4a

Observation ad0d7d1a-0d36-4340-9651-9a858d447735 · outbound

This paper cites Beyond Classification: Financial Reasoning in State-of-the-Art Language Models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Beyond Classification: Financial Reasoning in State-of-the-Art Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:15.739015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:15.739015Z digest=sha256:639fc8a4747eebf7d8a36fead7b0632397d2528217a48c68695e800f2c618109

Observation 3201edf1-4f3e-4d99-a74a-b41de35e9f1c · outbound

This paper cites Multi-step reasoning in Korean and the emergent mirage.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Multi-step reasoning in Korean and the emergent mirage

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.945080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:15.828572Z digest=sha256:887b416c05b55a1cfce1888de98b25a02afa93c5c8b456afa44ba91a1f087875

Observation 624cdaba-8ce7-4ba1-b8f4-cfe3cd287df1 · outbound

This paper cites KMMLU: Measuring massive multitask language understanding in Korean.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KMMLU: Measuring massive multitask language understanding in Korean

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.774825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:15.911316Z digest=sha256:703a08378caaa76798569552782a95d57a7b3993784c3a747fea73f908a32182

Observation 062acff7-3666-48da-83a3-d3376048856b · outbound

This paper cites HAE-RAE bench: Evaluation of Korean knowledge in language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation HAE-RAE bench: Evaluation of Korean knowledge in language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.551437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.003057Z digest=sha256:30a3fd91a33d086b9bb49ec5e2afb659adb0bbe9a4c6cab8160b6475b145dd83

Observation 51cdb724-54f3-47c5-9fc7-280e106df693 · outbound

This paper cites Challenging BIG-bench tasks and whether chain-of-thought can solve them.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Challenging BIG-bench tasks and whether chain-of-thought can solve them

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.366551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.099651Z digest=sha256:ce9f182ea92ba2b3cd30be3f4658addd6cc29ae25fbec15823a7c48ba9e51d71

Observation f7a0cb8e-86be-4b03-ac03-4c9eb57a57ad · outbound

This paper cites MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.178310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.178310Z digest=sha256:cc61ddda81eb4d93e10fe462856aa994db834959c8444b1119bc20669eebcd4b

Observation 432bf232-706d-4036-b311-207966c0a5a0 · outbound

This paper cites CommonsenseQA: A question answering challenge targeting commonsense knowledge.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CommonsenseQA: A question answering challenge targeting commonsense knowledge

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.285247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.285247Z digest=sha256:07475525e08a5c2246d176a8bea0ec51984b288d4fc838cba9ac73258c71e43d

Observation 00d892b9-90c2-4f7b-a54b-7846b57cd4c5 · outbound

This paper cites Gemma 3 Technical Report.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Gemma 3 Technical Report

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.366173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.366173Z digest=sha256:a65cb0b471a74460588a8d2e47eb28e2e0d1414fdd0800f9b933f5e3f60bb046

Observation 4b644f29-8f98-4946-a501-c1d937aec4bd · outbound

This paper cites Benchmark suites instead of leaderboards for evaluating AI fairness.Patterns, 5(11):101080, 2024.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Benchmark suites instead of leaderboards for evaluating AI fairness.Patterns, 5(11):101080, 2024

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.206356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.462836Z digest=sha256:93093d17b8681c7139bbcc9d46ad695f73f6999f274c76d511c57df6f2a81fed

Observation 6098c32e-3b67-48eb-92a2-d20244f6f03a · outbound

This paper cites SeaEval for multilingual foundation models: From cross-lingual alignment to cultural reasoning.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation SeaEval for multilingual foundation models: From cross-lingual alignment to cultural reasoning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.993735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.586923Z digest=sha256:e201f8d511fc5d76afe0c78645469c19991a088ae00bb18418ba150b504022a7

Observation 1fd0d21e-54f8-4e60-83f6-9b48a90b3a76 · outbound

This paper cites KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.759287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.759287Z digest=sha256:8f4f2bd0ebad38d4bfa567e6be7850d5019cf425cbc9231504054832bec25499

Observation 49db165a-1c43-4554-97df-544abd4e6381 · outbound

This paper cites Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.840752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.857171Z digest=sha256:18c32bca0285367f04f4240ca8c9bce9440e701b21b53963dce5985701729a5f

Observation af13bd18-0319-4209-a00e-594f1aa283f0 · outbound

This paper cites MMLU-Pro: A more robust and challenging multi-task language understanding benchmark.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MMLU-Pro: A more robust and challenging multi-task language understanding benchmark

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.656148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.951540Z digest=sha256:08903633d327ac1aa0ee256e8978f5d59c5dd4b5fe91bb3311da7607f037dbc4

Observation f44943f3-cb76-4bfa-aeeb-bf86bf08a589 · outbound

This paper cites Toward an Evaluation Science for Generative AI Systems.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Toward an Evaluation Science for Generative AI Systems

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.035975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.035975Z digest=sha256:024517f16a493eb1e7c6e384932f11bf431425a9268569144edbf26d2fb957ea

Observation a3199245-1f8c-4d78-958b-bed6e0247619 · outbound

This paper cites Qwen3 Technical Report.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Qwen3 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.152295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.152295Z digest=sha256:ed00305c9707c01641ffedfb389a64361dad6e9fb7b2e5c85d62e4c4b731e1d1

Observation e2d4076c-8373-42b7-bbb3-a99321d88f7a · outbound

This paper cites Qwen2.5 Technical Report.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Qwen2.5 Technical Report

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.279546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.279546Z digest=sha256:bb638ac33322b2f1015e26485087f88077656f84a04735d878151b3a131ea76c

Observation 0a0588e0-0c37-4bd2-95fb-f1c7e168d36e · outbound

This paper cites Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.369531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.369531Z digest=sha256:a100d5d481d4c7e71944c2b0048e215e196234a6f0aadace65e23db2e46ca261

Observation bc172362-c9b7-4ef7-9b97-e23956a44a2b · outbound

This paper cites FLASK: Fine-grained language model eval- uation based on alignment skill sets.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation FLASK: Fine-grained language model eval- uation based on alignment skill sets

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.440683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.447929Z digest=sha256:3fbf56aa0b16bc937e30eca07e1daf4f9b690c984f9403f308792ab3bd22c532

Observation 7da61289-4ca5-42d3-86b8-83542a2efdc6 · outbound

This paper cites GeoM- LAMA: Geo-diverse commonsense probing on multilingual pre-trained language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation GeoM- LAMA: Geo-diverse commonsense probing on multilingual pre-trained language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.253520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.503711Z digest=sha256:39c98030efde5ecb91633525309de5fdbe42ce1129a0c3cdc3153f1c8c5194ba

Observation 725c2996-fa56-420c-b4f4-d17e36ae7af6 · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.614168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.614168Z digest=sha256:5a7ea67a5e4ddbf46299bf372fa68209f845fb26a99fb5ff9bd34dc3f15ecb51

Observation ad13bf2f-b95b-429a-96e4-7f7b91ef36e0 · outbound

This paper cites Maruf Hossain, Guang- Jie Ren, Kate Soule, Yifan Mai, and Yada Zhu.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Maruf Hossain, Guang- Jie Ren, Kate Soule, Yifan Mai, and Yada Zhu

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.085125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.692121Z digest=sha256:96437d643762c39200e27b5e887b108c9fcfc5f688b9b5c92d7780759f13d589

Observation 29a84389-3535-41e3-85c8-51ccf4318ed9 · outbound

This paper cites Task me anything.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Task me anything

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.975615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.749553Z digest=sha256:919133b190382f8e0df57f6ed3b6801d8c4ab7ec7e62e58756c0240b8ca2d335

Observation 86d268ff-e827-4c90-a2e4-463393160202 · outbound

This paper cites Users can interactively explore the overall data distribution they are interested in.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Users can interactively explore the overall data distribution they are interested in

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.847818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.836203Z digest=sha256:c8105499a41fb4f551002534ad987dab3faa3d158ca657b89b3cf0d53ecfdd8b

Observation e0de9c67-b6cb-47e7-acf9-528c7ec99ed7 · outbound

This paper cites By reviewing samples, users can verify whether the dataset matches their needs and explore datasets suitable for their purposes.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation By reviewing samples, users can verify whether the dataset matches their needs and explore datasets suitable for their purposes

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.739965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.905218Z digest=sha256:e0a4153fd1f3ef6f39f088ae2520f9fce11f8e1fd57fa5bc81ff791e85da00db

Observation 7f2adf41-922e-4c81-9f4c-6dc5c22c83cd · outbound

This paper cites Is the Earth flat?.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Is the Earth flat?

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.606763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.955891Z digest=sha256:6babe48631172d1ff6e7426ac4bfbb8e94426e3d4a7e4d3d1a6175eeb4e759a8

Observation 88c7fafc-7f12-49be-82bc-bc53c5530e79 · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.678445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.678445Z digest=sha256:f01fb884cf2edda371948f4c7387319400810db795b0df513bac85528e3d9014

Pith citing papers

No inbound Pith citation observations are available.