Pith. sign in

Paper Citation Record · LEDGER

Holistic Evaluation of Language Models

As of 23 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 100 inbound Pith citation observations for arXiv:2211.09110.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.09110 v2

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 121 of 121 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 100 of 389 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:08:24.813142Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved4
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch7

External citation measurements

16
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation fd2b9ddb-6b51-4c64-b546-8f994c8bb60e · outbound

This paper cites Language Models are Few-Shot Learners.

Holistic Evaluation of Language Models Language Models are Few-Shot Learners

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T10:09:19.266423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:0f20eb9ea534cb3dc68f9bec188eba032679d0da8810328a3cf6a4733c2b5c1c

Observation 4fc7afca-6d7e-46e3-af1d-835c2d1bc741 · outbound

This paper cites doi: 10.18653/v1/2021.acl-long.150.

Holistic Evaluation of Language Models doi: 10.18653/v1/2021.acl-long.150

Reference 2

Resolution
verified exact
doi, observed 2026-05-24T10:09:19.299324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:9996f408e7c83163970b20fda6ce37c804966da935254565cb6ebb99672861cd

Observation d1b20050-8ea6-4d4b-ba5b-c3abdc75b098 · outbound

This paper cites URLhttps://glottolog.org/accessed2021-08-08.

Holistic Evaluation of Language Models URLhttps://glottolog.org/accessed2021-08-08

Reference 3

Resolution
verified exact
doi, observed 2026-05-24T10:09:19.283766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:75ff6db5b8f521eefd2ec5f4b1edb0fd92a506fb7ff852e67d026a32a3f0501c

Observation 7b7f0655-19dc-4471-b9be-85b3bb247e18 · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

Holistic Evaluation of Language Models Measuring Coding Challenge Competence With APPS

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T10:09:19.289641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:ae5994753790ac360d7f1175b8bc8af079e2ae5394a0b51655fda6b7f9cf3059

Observation 8e60302d-56ca-4737-8cc1-67f5948ec5e3 · outbound

This paper cites In Christopher Hitchcock & Alan Hajek, edi- tors: Oxford Handbook of Probability and Philosophy , Oxford University Press, pp.

Holistic Evaluation of Language Models In Christopher Hitchcock & Alan Hajek, edi- tors: Oxford Handbook of Probability and Philosophy , Oxford University Press, pp

Reference 5

Resolution
malformed identifier
arxiv_id, observed 2026-05-24T10:09:19.254183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:87fdcecd0f5c0561290eea418297ced8dd2d4aafd51beec4e8c272bf9631cd78

Observation 333e3eee-7498-4e14-a12c-24d304c33015 · outbound

This paper cites Cognition , year =.

Holistic Evaluation of Language Models Cognition , year =

Reference 6

Resolution
metadata mismatch
doi, observed 2026-05-24T10:09:19.273599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:c7d1df0a2696c0a9544fa7166736a71de0bb36a1e7b030cc1c82bfac11ab3f28

Observation 655b68ab-0010-4a0d-9fbf-3528748e4909 · outbound

This paper cites The Natural Language Decathlon: Multitask Learning as Question Answering.

Holistic Evaluation of Language Models The Natural Language Decathlon: Multitask Learning as Question Answering

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T10:09:19.279889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:a75faade5e9b2bc1280188190e90d0334f1d6cf9264adc4800ee40f4a6cedebe

Observation c7150b94-318f-4e30-8f44-5d369f022534 · outbound

This paper cites Red Teaming Language Models with Language Models.

Holistic Evaluation of Language Models Red Teaming Language Models with Language Models

Reference 8

Resolution
malformed identifier
local_arxiv, observed 2026-05-24T10:09:19.693997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:4556861b986958b8f60d9191e5a1a7e383815e6b760cf205345a2d6ffaf393d3

Observation c7b09df6-a83c-4bda-affe-0aebbf696222 · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

Holistic Evaluation of Language Models Measuring and Narrowing the Compositionality Gap in Language Models

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T10:09:19.294997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:a73c5d5316ae07ff8b71305b239effa82f5390a875d3fe90aa7d12a24a424efa

Observation 02724bfe-2e4b-42b2-beaf-c73880838e9c · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Holistic Evaluation of Language Models BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T10:09:19.247836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:a44be670ba881de42f47af3ec6ef77a2afcaf2172aa6686812d88aa390951ad4

Observation bbe34219-0bfc-444e-9d46-7e8bfdb9d9c2 · outbound

This paper cites Yes” or “No.

Holistic Evaluation of Language Models Yes” or “No

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-24T10:09:19.260372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:979ca91767d1e7037bbdd5f1dc62ba0c5bf3ee9ff83ab1b0496bae1859a98c11

Observation 73b5c53d-ac5d-40d7-80f5-ae782b89d9e2 · outbound

This paper cites an unresolved cited work.

Holistic Evaluation of Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:19:19.794025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:3f4365fa01770ee0dbbda7672b67a710bf55044e9dbe34b8593615bf327fde71

Observation 99063e66-2f6c-4e86-9ca2-88dc4e348cd1 · outbound

This paper cites an unresolved cited work.

Holistic Evaluation of Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:19:19.797403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:08955648b8a327f40c7de03246a613ce49791b1dae8eb7f15c63ca4be6c7de02

Observation d10ca844-e51b-418c-b639-d9e2b9f95339 · outbound

This paper cites Grandfather.

Holistic Evaluation of Language Models Grandfather

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T10:19:19.813945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:f0140a131e71999926833ba3b9d72708c1831cf2add36b21362ae59f8eb99aa1

Observation 85ae1069-f0e0-4dac-9230-449b76a51acf · outbound

This paper cites (2017), which derives its list form Greenwald et al.

Holistic Evaluation of Language Models (2017), which derives its list form Greenwald et al

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T10:19:19.805151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:dadf6bedabace08008e0ed5f43747244099091e0eb895563679843b35d0999dc

Observation edd232ba-c256-498d-8d6d-8b31e9d535ce · outbound

This paper cites (2018), which derives its list form Chalabi & Flowers (2017).

Holistic Evaluation of Language Models (2018), which derives its list form Chalabi & Flowers (2017)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T10:19:19.811250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:bec32cc6c63394b2fd20712cf57f7488cf108b84585665d715a243db07633b8e

Observation 20e4603d-6d17-4152-bf89-71769e843b82 · outbound

This paper cites It came from down here.

Holistic Evaluation of Language Models It came from down here

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T10:19:19.802573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:dc0f94e3b47c1c7d79f9ae34c7b9357422e32db3d3a93f64acb7cf43b92b272e

Observation 740f62a9-1862-4a24-bf73-b9fb2a035933 · outbound

This paper cites an unresolved cited work.

Holistic Evaluation of Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:19:19.800328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:e7d57a99cce1faef6d2ec85c5976fbcee331732badf67072c67600f4fa45e8cb

Observation ce64fb07-b7be-4c50-8c71-8c8963c0a340 · outbound

This paper cites an unresolved cited work.

Holistic Evaluation of Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:19:19.807714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:a4766f102c6868ac7b8d9382ea161895dbd55617558fc71576232af8adca44fe

Observation 37e2122b-5091-448f-890e-09361fe234a0 · outbound

This paper cites The capital of France is __.

Holistic Evaluation of Language Models The capital of France is __

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T10:19:19.791793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:b4a074e28167a02424f76ea5338c46a3598ceb7e4bb83c541fff391c8e20bc63

Observation 95af2df3-d032-4c58-b63e-9a27e2ff1c79 · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Holistic Evaluation of Language Models BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T10:09:19.667240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:1a79de77588d522453d3b0d7a145baead7c68b6ec86d2bc2f598cfae1a5e13e8

Pith citing papers

Observation 2a2e31f5-aeec-4826-b78d-e5c493f93d40 · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Holistic Evaluation of Language Models

Reference 187

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T00:51:11.310415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:011857f799cb14bf0bc72c45a7ed1b990768bc5409322148b1ee2932eeba6f46

Observation 21df8cda-88d8-46fd-8b46-d991c83a6b65 · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Holistic Evaluation of Language Models

Reference 266

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T00:51:11.617142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:53aafde083c25e1d3e2fdd084eed452fbbd141205541f53024e497adeff185dd

Observation b857d7ae-e7a1-4bad-b72b-3dc02cf82a60 · inbound

BloombergGPT: A Large Language Model for Finance cites this paper.

BloombergGPT: A Large Language Model for Finance Holistic Evaluation of Language Models

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T23:19:46.406429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T23:19:46.231145Z digest=sha256:7d1b11b73ec1077c4106ea9edda59bd0bf8108c347e6cdc553d3878525a56732

Observation 71fd7756-0d77-4e79-98c8-079e786055e0 · inbound

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models cites this paper.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Holistic Evaluation of Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.393937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:96d4fb3037023f9a2e0db4fac907bcc11f97c26d4d062f552482563691c7a260

Observation 109acafb-4654-4b97-ade5-b7d1d0f50f30 · inbound

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting cites this paper.

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting Holistic Evaluation of Language Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T12:01:19.833619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:580bf4fdb0f85b02c4f2fa5ecefed3d70b800c263e03e0fedab6c9ca978805f3

Observation 284311c4-f8f7-4c9e-9863-d59a231203bf · inbound

StarCoder: may the source be with you! cites this paper.

StarCoder: may the source be with you! Holistic Evaluation of Language Models

Reference 278

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:33:01.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T23:32:59.517389Z digest=sha256:a96270d8b2fb2b89049d93a37dabc04d735d48177a43a89bc7094cdd1266a033

Observation 2fac3c34-f5b5-44a0-95d9-1cfb280ef061 · inbound

Towards Expert-Level Medical Question Answering with Large Language Models cites this paper.

Towards Expert-Level Medical Question Answering with Large Language Models Holistic Evaluation of Language Models

Reference 87

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T04:32:33.596042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-24T04:32:33.271634Z digest=sha256:d02ebf1de41e422adc98b34eb8d544f5019b3163f5dcaa693ef09db974845e8a

Observation ccaf4683-caab-4354-85e0-1b4d6528f6c1 · inbound

PaLM 2 Technical Report cites this paper.

PaLM 2 Technical Report Holistic Evaluation of Language Models

Reference 91

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T11:59:27.187433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T11:59:25.813128Z digest=sha256:e489b14610535a8a5790caef12aed4a9708ecc6c6bd9720697b4ae571a851c99

Observation 3114b01c-f1c7-4dac-b5ad-6c9c791af106 · inbound

QLoRA: Efficient Finetuning of Quantized LLMs cites this paper.

QLoRA: Efficient Finetuning of Quantized LLMs Holistic Evaluation of Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:29:53.548060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T13:29:53.345251Z digest=sha256:c459739066f580b88324c7b3d4acaaa9bdc78e73f547dd60a0f4970e960cb116

Observation c50b9ac0-713b-4088-a30b-8a22fa424d79 · inbound

Scaling Data-Constrained Language Models cites this paper.

Scaling Data-Constrained Language Models Holistic Evaluation of Language Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:35:21.244425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T01:35:21.150772Z digest=sha256:ac93157b036ccce941eac8f5c2b5ee34e90215577e6e7e22fc8ffd8132f4bfa4

Observation 4c4b2457-6247-403d-a0b9-40662455db54 · inbound

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena cites this paper.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Holistic Evaluation of Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:52:59.095035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:b789ad931a65b3d3cf98274d1e09b1cc7c1b4977180a4bc24df5ae25d110cf9e

Observation e7f06c6e-425d-4089-b4da-5e359e40c414 · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Holistic Evaluation of Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.427058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:8a1bc8a22757fd8b9d24ee3570d039557c1089f29e61ceee3cc5b733d83ef43f

Observation 4db573e6-6781-430a-b62c-0d42808378ed · inbound

Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment cites this paper.

Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment Holistic Evaluation of Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:30:44.671076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T22:30:44.520703Z digest=sha256:fec9d5b412617dfa73e882962a70b9ee491783f76e337bc00790c5bcd2fe1c5e

Observation 2276c250-7782-4a1b-9da1-8fc07679cd7c · inbound

GAIA: a benchmark for General AI Assistants cites this paper.

GAIA: a benchmark for General AI Assistants Holistic Evaluation of Language Models

Reference 115

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T15:46:03.400665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T15:46:03.247029Z digest=sha256:a70ef3591bd51589e172953a69471f7798de80ddaf4f8a7adef6319118251db5

Observation f3811a93-4e13-4ad1-945e-4db5cd664d97 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Holistic Evaluation of Language Models

Reference 212

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:46:09.906965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:4f13d057fd99dde3cbce81c26e9a08b93a888241e86472c812ea36deb254cbe2

Observation dacd4082-fada-42c9-bcd6-eccb45ff3db4 · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Holistic Evaluation of Language Models

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:17:08.428332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:1356f34c1a385d1d0716abc0a947cd96a24835ede22d96b5360852b5959e7c11

Observation 24054ef2-9599-4b38-be4d-b42b0b32f8c9 · inbound

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive cites this paper.

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive Holistic Evaluation of Language Models

Reference 121

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T23:04:44.477071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-17T23:04:44.287660Z digest=sha256:13e3a9ec9fce3b2d15b6ea5a6903f0517f920cde86c199b4e53e8e6e1b53c9cd

Observation 8a3635d8-5b20-4c17-9a12-330413744d0b · inbound

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark cites this paper.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Holistic Evaluation of Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:08.467354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:dc7c538e7ea9cba48146b259415a334f3b64a9a06d477f8c806eb9ef67f0a923

Observation ca8c94d0-bb92-4d0e-ba17-fd0a9cad37e2 · inbound

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline cites this paper.

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Holistic Evaluation of Language Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T00:03:38.982197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:90f84506ca4948dd2eca5acb9af8a72680169bc3b0fc4b30b38a96ac284a12fe

Observation d4b8d1bf-352d-402b-b588-85accc1e953a · inbound

Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models cites this paper.

Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models Holistic Evaluation of Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:18:31.890032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T22:15:52.638622Z digest=sha256:54322a697a702ad83c25b5477219c2c46b047982d77cca10e72d1cad24cdbaf4

Observation 4ceacae8-a589-4f00-a092-b3fe1463346b · inbound

The Limited Impact of Medical Adaptation of Large Language and Vision-Language Models cites this paper.

The Limited Impact of Medical Adaptation of Large Language and Vision-Language Models Holistic Evaluation of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T21:16:35.248965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:16:35.248965Z digest=sha256:a3599259026b23b897bb013c5c10c8511b0de603f3515603928ddb40677232e8

Observation f666593b-a629-4ee6-9b84-18f6c44e7f36 · inbound

Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents cites this paper.

Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents Holistic Evaluation of Language Models

Reference 149

Resolution
unresolved
no resolver link, observed 2026-08-12T20:36:02.007205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:36:02.007205Z digest=sha256:7c499aa8e3d1d18bf40c79b1bf1b051c6f65771f74610aab1bb875b17866cd6f

Observation 1d530560-b6e5-4118-b30a-9f1c0fc96e01 · inbound

Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems cites this paper.

Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems Holistic Evaluation of Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:31.073347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:31.073347Z digest=sha256:4f3caf3fcc7c3a1621c563a2d0de145c99c53e2c7bbe55e6c6aeb97b9c31c140

Observation a8a115a4-1bc7-414d-8a1c-72726c832982 · inbound

IntentGPT: Few-shot Intent Discovery with Large Language Models cites this paper.

IntentGPT: Few-shot Intent Discovery with Large Language Models Holistic Evaluation of Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:31:34.368539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:31:34.368539Z digest=sha256:570219ab500996050b2e9ac3901c70bd3b76977c9d98ab65ec3dac4dcf6e5a1b

Observation 854a1e9f-aa4b-4011-9d9e-4a05206611dc · inbound

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs cites this paper.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Holistic Evaluation of Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.655265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.655265Z digest=sha256:02104baa96d8c690777089b3d234737b0e3352bf203f81f3854ceb9e8db5dae4

Observation 6effd4b1-d362-43ff-9d01-695d8a0137ea · inbound

Dimensions of Generative AI Evaluation Design cites this paper.

Dimensions of Generative AI Evaluation Design Holistic Evaluation of Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:21.919619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:15:21.919619Z digest=sha256:a1a481c59acaac924128860aefa528dbbfc09fe724981bac2c99eb87b0a17b43

Observation 8b535e1e-46e6-4529-b7dc-dfb5862ecb3a · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models Holistic Evaluation of Language Models

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.647383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.647383Z digest=sha256:f4ba82b7575c7d143ccec6e4d37481e63bc96e5e4a77d9001508191b2b5bcd49

Observation c4e1c07a-6472-4e33-8bce-b595859eb664 · inbound

GPAI Evaluations Standards Taskforce: Towards Effective AI Governance cites this paper.

GPAI Evaluations Standards Taskforce: Towards Effective AI Governance Holistic Evaluation of Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:54:33.121732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:54:33.121732Z digest=sha256:eb10f1e039996bf3566413c9e78d56a191a8b82c87cff81e12a70ce009c3229b

Observation 27f63275-d104-41ff-9613-56c2c419aae4 · inbound

Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in Medicine cites this paper.

Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in Medicine Holistic Evaluation of Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T16:56:44.738374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:56:44.738374Z digest=sha256:267777dabfbf51ff0e61f2d70c7b0f5928d5dcabe72161bd0cc362ee7a0c2440

Observation 45e17965-2ffe-431c-bd3b-2ba97c179f1d · inbound

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models cites this paper.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Holistic Evaluation of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.226325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.226325Z digest=sha256:2e557073a68a48510aa74c10ed8871318cea93fa1d5697f0584a793fc7b08eb6

Observation dbf6de75-1d22-4c99-acec-3c2b57599080 · inbound

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set cites this paper.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Holistic Evaluation of Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.413062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.413062Z digest=sha256:75de5f6fd00cbbd6785451a67ab316f91c714b5021893084f68aff0be081e3fc

Observation 38462fc9-d789-48ce-8a39-a724681605a7 · inbound

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity cites this paper.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Holistic Evaluation of Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T13:27:06.690101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:27:06.690101Z digest=sha256:08d2f740e7253a6feb1f55a0c566b3c9bc624464ef0255b053d32f1ff6084e31

Observation e4b60683-06f7-44b2-9629-5734beb4a3c3 · inbound

Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks cites this paper.

Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks Holistic Evaluation of Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:28:14.337637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:28:14.337637Z digest=sha256:b1da7a651591f70cff1dd2d1e470fc0e9d3ea2165dd95cdeeb06e6142b66a38f

Observation 7ffce77d-4554-48c2-b7a0-00122bd02305 · inbound

Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach cites this paper.

Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach Holistic Evaluation of Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:21:33.895762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:21:33.895762Z digest=sha256:022cad80ab6f6cb4fc2d09bb56eb5d6ebf2eb949ed702bcc2993886efc0ee238

Observation 4e7290a1-7f3f-4dad-81a5-8aaeb574c22d · inbound

Enhancing Zero-shot Chain of Thought Prompting via Uncertainty-Guided Strategy Selection cites this paper.

Enhancing Zero-shot Chain of Thought Prompting via Uncertainty-Guided Strategy Selection Holistic Evaluation of Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:56.136541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:56.136541Z digest=sha256:c66f8394c80955d535f17b2e233cf4c690468dac78d078f261e9050765a462d0

Observation 63c82884-8e57-4f5a-8d98-b09919060252 · inbound

Rank It, Then Ask It: Input Reranking for Maximizing the Performance of LLMs on Symmetric Tasks cites this paper.

Rank It, Then Ask It: Input Reranking for Maximizing the Performance of LLMs on Symmetric Tasks Holistic Evaluation of Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T05:21:55.117620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:21:55.117620Z digest=sha256:37f25d1c9e42b4d80f3fe4ebc909d51e2ef16c913c913fde3cb67478f0c6fb87

Observation 33d99f40-47cc-4c61-aea0-6812515a4d73 · inbound

Unveiling Performance Challenges of Large Language Models in Low-Resource Healthcare: A Demographic Fairness Perspective cites this paper.

Unveiling Performance Challenges of Large Language Models in Low-Resource Healthcare: A Demographic Fairness Perspective Holistic Evaluation of Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T05:18:55.812296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:18:55.812296Z digest=sha256:33e181913a0eaf55ea8ba8417b1796db03e15188cf6c19adf782f05b0816e7e5

Observation 627e066d-cf78-46da-beef-dbc4afd38e81 · inbound

Best Practices for Large Language Models in Radiology cites this paper.

Best Practices for Large Language Models in Radiology Holistic Evaluation of Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:59.502032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:35:59.502032Z digest=sha256:f895912dd7bff9136d7b0884c2ec27bfa34f38fe446620d82cbf4124b59fd7c8

Observation ffd2df8d-b59f-4695-9cc1-eb5b565deece · inbound

Addressing Data Leakage in HumanEval Using Combinatorial Test Design cites this paper.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Holistic Evaluation of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.483043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.483043Z digest=sha256:9fc427d560ce4e029d5aa167dacf5ae32b0893d189ebee797dc3a8678db7f7c6

Observation dc408e55-2172-4b16-ae57-05e086bfd276 · inbound

CopyrightShield: Enhancing Diffusion Model Security against Copyright Infringement Attacks cites this paper.

CopyrightShield: Enhancing Diffusion Model Security against Copyright Infringement Attacks Holistic Evaluation of Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:02.813497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:02.813497Z digest=sha256:f8276ad5ce815a7005283c1da817a5b417076ceeb31b46c253a6b84bfadbd736

Observation 396228b2-8607-415e-a327-5bec2dd767af · inbound

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization cites this paper.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Holistic Evaluation of Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.303089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.303089Z digest=sha256:8935fdba3b97ec6393449486fefb7b4d49de6a7d643ac70009ca53778eab2700

Observation 496bbc29-aa5d-4013-9cc9-e5810ee50300 · inbound

C$^2$LEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation cites this paper.

C$^2$LEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation Holistic Evaluation of Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T21:10:35.151341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:10:35.151341Z digest=sha256:e877f893f1d5a39d5633a68eedd6ed4a0d694cc7f34f945278b690d137b82633

Observation c29116bd-d715-461e-bc48-61ca2bbfff9a · inbound

Code LLMs: A Taxonomy-based Survey cites this paper.

Code LLMs: A Taxonomy-based Survey Holistic Evaluation of Language Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:57.173433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:57.173433Z digest=sha256:7c13457768182fffdad1f55d5f4c13f1af4cf61a6ffb08582abb7796459a2682

Observation bd859e7b-4303-43e9-85be-3d014470063d · inbound

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model cites this paper.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Holistic Evaluation of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.759469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.759469Z digest=sha256:39c6cdd63e716fe258dcd7c8919279b1ad14728ca38316f7c907f064ec4ae99c

Observation 84169d9f-473d-4b0e-ae81-120693f2d837 · inbound

ReFF: Reinforcing Format Faithfulness in Language Models across Varied Tasks cites this paper.

ReFF: Reinforcing Format Faithfulness in Language Models across Varied Tasks Holistic Evaluation of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:21.846142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:19:21.846142Z digest=sha256:0a1281b1d59e5ea4d25500df2b1f62ed2eaafb4dbed38c5c4095990b9023a030

Observation 3f8c4ce8-bd46-455b-b93b-5700c83e709a · inbound

Biased or Flawed? Mitigating Stereotypes in Generative Language Models by Addressing Task-Specific Flaws cites this paper.

Biased or Flawed? Mitigating Stereotypes in Generative Language Models by Addressing Task-Specific Flaws Holistic Evaluation of Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:33.258973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:02:33.258973Z digest=sha256:687dd4e235f3a98cbb884932930a9b084a7dbb1cd9453dc70880bdfd19ecf9bb

Observation add45c71-8f27-4ea8-87d5-bd84b9256717 · inbound

PickLLM: Context-Aware RL-Assisted Large Language Model Routing cites this paper.

PickLLM: Context-Aware RL-Assisted Large Language Model Routing Holistic Evaluation of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:29:02.773849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:29:02.773849Z digest=sha256:11b818212e520848524e9ea464e9a862d01364360d55ffecd1c592f2432d3fc1

Observation bcdbf6ba-05ba-469f-aa78-396ca4958c90 · inbound

Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation cites this paper.

Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation Holistic Evaluation of Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:58:26.892141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:58:26.892141Z digest=sha256:138203d740cda0ffb7866a5be6a6b80aa706ac78a699427b9561301c7456c10b

Observation 9f689af3-fa83-48a2-b8ad-cb04e59ac950 · inbound

Sliding Windows Are Not the End: Exploring Full Ranking with Long-Context Large Language Models cites this paper.

Sliding Windows Are Not the End: Exploring Full Ranking with Long-Context Large Language Models Holistic Evaluation of Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:09:29.199485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:09:29.199485Z digest=sha256:737684b5476f5c8f193a31fd398a180c8692c892eb30914bc8dc1d28df14f2bb

Observation aaf64c9c-8a63-44e6-abab-2941bf075d34 · inbound

INFELM: In-depth Fairness Evaluation of Large Text-To-Image Models cites this paper.

INFELM: In-depth Fairness Evaluation of Large Text-To-Image Models Holistic Evaluation of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:48:02.339502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:48:02.339502Z digest=sha256:27ec8098b773be24e12396d2dd12eda7c6a71f613d5c800b11cc8ac65de0e59f

Observation f7e3189a-48ed-4eda-8eb0-f730cbd65023 · inbound

LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases cites this paper.

LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases Holistic Evaluation of Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T21:56:29.802738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:56:29.802738Z digest=sha256:9fd2dba513c86d13473d03e95afa95fd60e16011c49c11b3ca4d521c05020507

Observation 24d4414f-df42-44b7-8b4f-fe720994635c · inbound

A Generative AI-driven Metadata Modelling Approach cites this paper.

A Generative AI-driven Metadata Modelling Approach Holistic Evaluation of Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T16:32:07.811499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:32:07.811499Z digest=sha256:352d90103b7ecab304ecf4dd305e43791e8c1b27bbf380345b2ac75b2b93b336

Observation 11554f1f-f4e1-40dc-bc36-013fbacd976d · inbound

A Survey on Large Language Models with some Insights on their Capabilities and Limitations cites this paper.

A Survey on Large Language Models with some Insights on their Capabilities and Limitations Holistic Evaluation of Language Models

Reference 190

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:56.116010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:17:56.116010Z digest=sha256:688c9d90b5f2b27e07ff05d574b8e7456107b840d6c97ead415fb4fa04badff4

Observation 9d6439d2-e77b-4d61-a9ac-b786c99d9dec · inbound

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values cites this paper.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Holistic Evaluation of Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.736259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.736259Z digest=sha256:0e0efce323cbad4403dfab20ae5e7946be641a19cad698df2741075ed5f69692

Observation 32469f6b-cb2d-49f3-b768-e031bb9cdc87 · inbound

Addressing the sustainable AI trilemma: a case study on LLM agents and RAG cites this paper.

Addressing the sustainable AI trilemma: a case study on LLM agents and RAG Holistic Evaluation of Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:34.841850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:34:34.841850Z digest=sha256:9faaa605a7ca28e665f91f013a55a3467f55e2bf6d3198738b20544803ab7a02

Observation 9f47109e-10b6-477e-b1cf-8a0cc5bbacf4 · inbound

Enhancing LLMs for Governance with Human Oversight: Evaluating and Aligning LLMs on Expert Classification of Climate Misinformation for Detecting False or Misleading Claims about Climate Change cites this paper.

Enhancing LLMs for Governance with Human Oversight: Evaluating and Aligning LLMs on Expert Classification of Climate Misinformation for Detecting False or Misleading Claims about Climate Change Holistic Evaluation of Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T15:39:33.230512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:39:33.230512Z digest=sha256:310493337d2dc95d35ce942f501fd28ed38f719fdf1b6fd2b4e7b6a0e0b7dcdf

Observation 5d7cc790-2d6c-4439-912f-fb955a6bec0b · inbound

RankFlow: A Multi-Role Collaborative Reranking Workflow Utilizing Large Language Models cites this paper.

RankFlow: A Multi-Role Collaborative Reranking Workflow Utilizing Large Language Models Holistic Evaluation of Language Models

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T04:37:31.933903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T04:36:32.897387Z digest=sha256:e680725475deb19ec0517f99b64ca1777f23e6e7722bbd7e7de888b4009b8a08

Observation 06999197-03dc-4598-b2eb-405ce3fdcf68 · inbound

Language Models Prefer What They Know: Relative Confidence Estimation via Confidence Preferences cites this paper.

Language Models Prefer What They Know: Relative Confidence Estimation via Confidence Preferences Holistic Evaluation of Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T16:34:01.122641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:34:01.122641Z digest=sha256:50038f76b25f3ce35d6263d4f07b435a37ba67c26ff39e02b14866fa52a0b420

Observation e6887e6a-0877-46d1-8f8a-f439804b5877 · inbound

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing cites this paper.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Holistic Evaluation of Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.913059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.913059Z digest=sha256:3c7adde32e0c18bd849fc5aa1c4fde929c077416a5e6519fb464ddbf3727a544

Observation 706c8388-aaf6-4f52-ad7d-565da31aaef2 · inbound

Powering LLM Regulation through Data: Bridging the Gap from Compute Thresholds to Customer Experiences cites this paper.

Powering LLM Regulation through Data: Bridging the Gap from Compute Thresholds to Customer Experiences Holistic Evaluation of Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:51:52.861137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:51:52.861137Z digest=sha256:d40ac156b106313ed72af9fd45789a7fa9b7175168d741b1cefc2ac8e546ffc2

Observation 651ef3b7-6b2e-4988-ad39-5486161d1280 · inbound

MultiQ&A: An Analysis in Measuring Robustness via Automated Crowdsourcing of Question Perturbations and Answers cites this paper.

MultiQ&A: An Analysis in Measuring Robustness via Automated Crowdsourcing of Question Perturbations and Answers Holistic Evaluation of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T01:05:23.072022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:05:23.072022Z digest=sha256:f9cc2eb02366fb292b03c07303bc5bc46343bc85f23b7e1b8688616b37b23b00

Observation f7e73f1a-be3b-4574-869d-ebff3a8d7aaa · inbound

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation cites this paper.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Holistic Evaluation of Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.036681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.036681Z digest=sha256:cef13d3c0bdd2f09893fdf6d1bcf3c0fc5be7f0356d93975f76c5ef91ae8e582

Observation eaeada13-0cc4-4762-9209-69f34cdc99f6 · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective Holistic Evaluation of Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.426227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.426227Z digest=sha256:7a9dee46f6004d50e52daf7fab5254c9e24651378d929a6feeba91639d1ae891

Observation e1aac824-a3ea-4db0-b811-de8ebc76936e · inbound

Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey cites this paper.

Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey Holistic Evaluation of Language Models

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:25.382826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:25.382826Z digest=sha256:62a79ffc47ac86f9d55b6531be0be723363c1b9ff832d5abb74de8b53bddb08c

Observation ce47f047-a3fb-4c34-a689-8de137f31e16 · inbound

White Hat Search Engine Optimization using Large Language Models cites this paper.

White Hat Search Engine Optimization using Large Language Models Holistic Evaluation of Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T13:10:56.377693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:10:56.377693Z digest=sha256:54065d4b42a217e931eae0c91c339e1a1c53511d38cb0e17cc15ad8f138f8140

Observation 6ccb9e5d-77dc-4605-97b2-3ff4ae39264b · inbound

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories cites this paper.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Holistic Evaluation of Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.733056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.733056Z digest=sha256:5ab018662a0a96a516e967a9827c6308cffa51c5c42aef3762e7f56659482e2d

Observation 29d6a1a3-2b5b-4296-9a88-b5697a9071e9 · inbound

Salamandra Technical Report cites this paper.

Salamandra Technical Report Holistic Evaluation of Language Models

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-08T04:58:33.096413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:58:33.096413Z digest=sha256:1d6623f037190c7f9e926f35af476314624aab068e95f47f5625798210449630

Observation 08eb5c24-8535-4b84-8652-abfed5201ad3 · inbound

Measuring Diversity in Synthetic Datasets cites this paper.

Measuring Diversity in Synthetic Datasets Holistic Evaluation of Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:50.768611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:54:50.768611Z digest=sha256:545476fabcf3b60c811a21100bfaf3e19efc9cf0763dcfa274353c672db1047f

Observation 88eb5930-81a1-40fd-bdd6-eebf5bcb7c84 · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models Holistic Evaluation of Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.752311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.752311Z digest=sha256:6ac35e874c34a99c8ea8ba6ee1338ffb927a1734117dab20ba5b272d9a440e70

Observation 298370d4-fe6e-481a-ae59-4f4ed00a8fd4 · inbound

Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages cites this paper.

Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages Holistic Evaluation of Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T19:19:38.204950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:19:38.204950Z digest=sha256:cd64099fbfc7fa1eb51c813bb037544894b2dcc71d8dd16807071d718ac0affa

Observation 53288b3d-354a-46b5-91bd-77643a3ce62d · inbound

Instruction-Based Fine-tuning of Open-Source LLMs for Predicting Customer Purchase Behaviors cites this paper.

Instruction-Based Fine-tuning of Open-Source LLMs for Predicting Customer Purchase Behaviors Holistic Evaluation of Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T10:07:33.771199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:07:33.771199Z digest=sha256:a149d2fb93eeebf9316bbb236063ba8d0ab0aa878a0f3090143aec66094b088d

Observation 7004c954-9449-48ce-8b1e-ba9edb649fcd · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist Holistic Evaluation of Language Models

Reference 169

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T13:02:45.077542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:b3c5de16018feb3707eeffc7f5b8c02083c76035178e566bd66c395ba423b97b

Observation f0bcf5e2-1a4f-4abd-9852-fb73ae166fe1 · inbound

Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation cites this paper.

Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation Holistic Evaluation of Language Models

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T23:42:16.115559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T23:41:22.017848Z digest=sha256:ab54f679fda84b1ada31898dd8b62d9d06c5884fbe25f10698142dbb8b9be675

Observation d2fbaf20-ec65-4cb3-b12b-59153d23f2eb · inbound

Exploring the Potential for Large Language Models to Demonstrate Rational Probabilistic Beliefs cites this paper.

Exploring the Potential for Large Language Models to Demonstrate Rational Probabilistic Beliefs Holistic Evaluation of Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:08:24.813142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:08:24.813142Z digest=sha256:e0af68c646d5dbc24b05752fcf624de71601890cadbabd34296ce5ff1b5c79a0

Observation d328c780-5f44-4b1a-9af2-d96949328338 · inbound

aiXamine: Simplified LLM Safety and Security cites this paper.

aiXamine: Simplified LLM Safety and Security Holistic Evaluation of Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T11:39:10.918707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:39:10.918707Z digest=sha256:0d5502b8f4eea2070844e4a54900f7f828ea5d22b179a096955f16e48b4100cf

Observation a93f9640-c6c5-4599-81fc-45e6f427374e · inbound

DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain cites this paper.

DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain Holistic Evaluation of Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T12:04:24.264556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:04:24.264556Z digest=sha256:0d71292b154540c728612f9b294e84d7cf81c712dc2410292a22980e6b2abf62

Observation 17351bac-c664-4485-a0ae-67855d85dc35 · inbound

PRIMETIME : Limits of LLMs in Temporal Primitives cites this paper.

PRIMETIME : Limits of LLMs in Temporal Primitives Holistic Evaluation of Language Models

Reference 80

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T18:36:58.792994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T18:36:48.376877Z digest=sha256:10a0145c40517caa6ba62260ac1782120e5f16314b52619c17beea680b3c2fbc

Observation 4e61d8c8-5d00-4acf-be30-a673f89df2b3 · inbound

An Empirical Study of Evaluating Long-form Question Answering cites this paper.

An Empirical Study of Evaluating Long-form Question Answering Holistic Evaluation of Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.005837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.005837Z digest=sha256:2c6c3b659c435f35c6931e3095bccbf0a1bcf8ab6dc8f5390675100eb979da75

Observation 12198f07-bd16-4b89-93ea-2ab43594bb64 · inbound

Mind the Language Gap: Automated and Augmented Evaluation of Bias in LLMs for High- and Low-Resource Languages cites this paper.

Mind the Language Gap: Automated and Augmented Evaluation of Bias in LLMs for High- and Low-Resource Languages Holistic Evaluation of Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:54:28.508916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:54:28.508916Z digest=sha256:155fe5b5e3e6d90d5070c253dce799194647a47aa419a838683f8545f89a4683

Observation 0ce24d1a-970a-41ea-a6bb-6f5aa27cf42e · inbound

Computational Reasoning of Large Language Models cites this paper.

Computational Reasoning of Large Language Models Holistic Evaluation of Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:24.018750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:24:24.018750Z digest=sha256:dadc91d67404a808a2b0fed469ef47fa2449fddb60c89d87a5cbcf603b61c4c2

Observation 1ad8c59c-a0d4-4d0b-8ca2-f856ac3b0ed8 · inbound

UrbanPlanBench: A Comprehensive Urban Planning Benchmark for Evaluating Large Language Models cites this paper.

UrbanPlanBench: A Comprehensive Urban Planning Benchmark for Evaluating Large Language Models Holistic Evaluation of Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.339143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:00:03.339143Z digest=sha256:2b94c6cd10c81500b17005982644f7597fc0c2e1c4d392e486e5b88938161b76

Observation 831e7284-0cf6-45ee-9524-d6c5dd07aa13 · inbound

LLM Ethics Benchmark: A Three-Dimensional Assessment System for Evaluating Moral Reasoning in Large Language Models cites this paper.

LLM Ethics Benchmark: A Three-Dimensional Assessment System for Evaluating Moral Reasoning in Large Language Models Holistic Evaluation of Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T04:37:48.205650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:37:48.205650Z digest=sha256:cb62f3dfda83d80de87f44735c9497c0f7e226462a4294d626e88f324072972c

Observation 20c0818d-9ac5-454d-8572-b459a4629379 · inbound

am-ELO: A Stable Framework for Arena-based LLM Evaluation cites this paper.

am-ELO: A Stable Framework for Arena-based LLM Evaluation Holistic Evaluation of Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:41.842272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:58:41.842272Z digest=sha256:b78a66409b32a63a551c8deff91feead99df2740073b15f9847e059c85546cf0

Observation 063e2f76-2474-4232-912d-b6090fcdedf5 · inbound

AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions cites this paper.

AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions Holistic Evaluation of Language Models

Reference 128

Resolution
unresolved
no resolver link, observed 2026-08-15T23:27:30.298577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:27:30.298577Z digest=sha256:59e7ff1aafff991bd134010f8a1611382c53f36c0b3b08cd0427754bb495c152

Observation 3184cf16-b833-4fc4-8ec3-ba3e28d3e9f5 · inbound

Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs cites this paper.

Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs Holistic Evaluation of Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:23:39.355813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:23:39.355813Z digest=sha256:c911b3dee0ebf0722c028de46686f300b0c7e5b6613f53422523a0f5fe419d39

Observation 71bb4ed7-d858-437c-b4b4-a8eece49a1ff · inbound

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics cites this paper.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics Holistic Evaluation of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.739467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.739467Z digest=sha256:635005534478416c29c1de7e01d2d53a2b22eadb44df1f8059529d5845f18a39

Observation 53e20df4-e36d-4458-8fff-d1076f7747c7 · inbound

DynamicRAG: Leveraging Outputs of Large Language Model as Feedback for Dynamic Reranking in Retrieval-Augmented Generation cites this paper.

DynamicRAG: Leveraging Outputs of Large Language Model as Feedback for Dynamic Reranking in Retrieval-Augmented Generation Holistic Evaluation of Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:09.495968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:09.495968Z digest=sha256:b2508b007d759a1afefc82cdf5d49e72a9925e1da780b70efdd75b9aab0c4730

Observation 805b589e-0e5e-4e7c-9142-466157ef57ed · inbound

Comet: Accelerating Private Inference for Large Language Model by Predicting Activation Sparsity cites this paper.

Comet: Accelerating Private Inference for Large Language Model by Predicting Activation Sparsity Holistic Evaluation of Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T22:26:12.477942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:26:12.477942Z digest=sha256:41c051fcaff172e6aa05eee0821fd0d9c59e8d2099005ed37d3ad979ee4fb34f

Observation 0671a7e5-2dcd-46c1-80b4-dff2aade5ce8 · inbound

Large Language Models for Cancer Communication: Evaluating Linguistic Quality, Safety, and Accessibility in Generative AI cites this paper.

Large Language Models for Cancer Communication: Evaluating Linguistic Quality, Safety, and Accessibility in Generative AI Holistic Evaluation of Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:57.709230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:12:57.709230Z digest=sha256:8d31380b42ccbe80c3df49d0a4efa13b5551c8639e11b967fa3790f9d17293bc

Observation aef0beeb-d0b6-4e56-bc9a-aa0e8ae9a766 · inbound

Phare: A Safety Probe for Large Language Models cites this paper.

Phare: A Safety Probe for Large Language Models Holistic Evaluation of Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:22.855640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:22.855640Z digest=sha256:0ca8459ed4fea9b8328b4816a8a3d0798050b84b0f6087a1a9255c539ec67450

Observation 1e1b4266-b683-466f-8c1b-d68a8833ed8f · inbound

IQBench: How "Smart'' Are Vision-Language Models? A Study with Human IQ Tests cites this paper.

IQBench: How "Smart'' Are Vision-Language Models? A Study with Human IQ Tests Holistic Evaluation of Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:47:45.502670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:47:45.502670Z digest=sha256:c13806488041bb4d7ede970c4b6db3553991c7faf11221716799255622226da2

Observation 166db7f0-7f4d-4bf6-bdc2-f6e8ebf26969 · inbound

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation cites this paper.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Holistic Evaluation of Language Models

Reference 1966

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.120283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.120283Z digest=sha256:66f9b4ef708093187f0b032b485bd84aaf59fd2692f8988b48ef650c522f40b1

Observation 7185838f-adde-4da3-89d5-65093b10910a · inbound

Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers cites this paper.

Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers Holistic Evaluation of Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:14:57.484055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T15:13:28.927880Z digest=sha256:856d5e350ee822a1f58f2d39cf56ea7e5e69c96663bae223741b69f0d1d5b809

Observation 6694cabc-06d1-4ad7-a6bc-433369601a04 · inbound

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models cites this paper.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Holistic Evaluation of Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.215708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.215708Z digest=sha256:818841521cd41228732c23df187dd0c5c64a7724ce3b151dbe437b3e5e139fdc

Observation 3950c820-dd66-486e-bf5e-b65e403992db · inbound

LLM-KG-Bench 3.0: A Compass for SemanticTechnology Capabilities in the Ocean of LLMs cites this paper.

LLM-KG-Bench 3.0: A Compass for SemanticTechnology Capabilities in the Ocean of LLMs Holistic Evaluation of Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:22:29.929741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:22:29.929741Z digest=sha256:1758c73c840ea650759a310fe63d9baf623f7abf79b42bcc780e4dd534cc303d

Observation 0b099e41-9d2c-41c1-abba-424777f6191b · inbound

TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents cites this paper.

TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents Holistic Evaluation of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:59.931309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:18:59.931309Z digest=sha256:e86cc27fb38d813c5131eab4c07c3cc2110128e8092fb2e64a2717dc5301dde0

Observation 0f962621-e721-4d8a-8038-118360d4715b · inbound

Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling cites this paper.

Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling Holistic Evaluation of Language Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T14:21:39.613433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T14:21:18.355360Z digest=sha256:b9e1f8ee7de4ceb5d13344cd6ad82573834a82c02d7a2e36369d1c3414ec3f7b

Observation 0167f08b-7e13-4d3f-a7c7-7424af1f6053 · inbound

Veracity Bias and Beyond: Uncovering LLMs' Hidden Beliefs in Problem-Solving Reasoning cites this paper.

Veracity Bias and Beyond: Uncovering LLMs' Hidden Beliefs in Problem-Solving Reasoning Holistic Evaluation of Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:13.198498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:12:13.198498Z digest=sha256:6c347a387aec8912c63a7e92cce6151bc317b292ecba9c5e09040399d039aebf

Observation 18a23f07-3b2a-4d2d-82e2-b3b24dd28cf7 · inbound

How do Scaling Laws Apply to Knowledge Graph Engineering Tasks? The Impact of Model Size on Large Language Model Performance cites this paper.

How do Scaling Laws Apply to Knowledge Graph Engineering Tasks? The Impact of Model Size on Large Language Model Performance Holistic Evaluation of Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:00.526886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:00.526886Z digest=sha256:f57d3ab6f876bf58f1bdbdb90d7a4caab3c56a60d0dfa109ec5d6aa9bbca88ee

Observation 9e817d53-75a9-4742-972f-1035cb756a88 · inbound

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs cites this paper.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Holistic Evaluation of Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:22.022328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:22.022328Z digest=sha256:b1ad5f5e406fed4c7a414cca5beeda92826b6896a2057e6a273fcf0cb0618e06