Pith. sign in

Paper Citation Record · LEDGER

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

As of 13 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 8 inbound Pith citation observations for arXiv:2509.04013.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04013 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:31:02.319589Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:46:08.844007Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T23:45:07.998622Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 128e5b78-77f7-44b3-8ba5-540d8af3ab63 · outbound

This paper cites Program Synthesis with Large Language Models.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Program Synthesis with Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.079925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.079925Z digest=sha256:3a7f01a0e9d9d7e57c24531c3b3071774f3f27b0ae8a341990c00ed84106f476

Observation 032f5bc1-5851-425b-9f95-bae93273db73 · outbound

This paper cites MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.085778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.085778Z digest=sha256:d5a5bb0b483000f7ad482fe015e239a832917f99803d67724b4b7542a3c4b544

Observation 5e28eac2-5474-4fa9-9b64-922e1d3182a5 · outbound

This paper cites Bailey, N.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Bailey, N

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.281185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.090860Z digest=sha256:c6f2b6f34d3ffaa0a58831bc3e891832fb5bf8d0691be5cef82f68402d4e0bfa

Observation 766cc208-33f1-4b9c-adb1-e312d203b2bd · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:03.267003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.095785Z digest=sha256:9591d489577ff60022e8bc62a1807ac70eb3f05e2242cbf9d1ec8857154fd639

Observation a3b2f874-c6b9-4767-9bdd-7e6b6e164b3a · outbound

This paper cites Burnell et al.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Burnell et al

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.252544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.101619Z digest=sha256:1243707d801bafdc652d192cfe0a95b038e5a5dd4fd93d14232eb80f0e9f715f

Observation 5c5a3774-d159-438d-8764-c88fab8a9d99 · outbound

This paper cites Carterette, J.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Carterette, J

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.237912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.106197Z digest=sha256:3ffcdcb01f7672c0a707a17b591ead3a38a42cb25a47b44b27945f9d5da48f36

Observation 555fd79b-d022-455b-9f0d-cdfb09ccc4a0 · outbound

This paper cites Carterette, A.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Carterette, A

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.223737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.111083Z digest=sha256:f9c9f55c33654e515d5bd7a812b72d2851ac2a155defaab525bb5599bb4474f9

Observation 0952470b-d78b-4187-b02c-03e45a157225 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.115716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.115716Z digest=sha256:cb7cdc08a1dac9614a21181a5e5ab726cddad78002ae1faf572361e3e6ac4058

Observation 73abcfd4-d6c3-4275-beb0-e0e23c633649 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:03.209510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.120756Z digest=sha256:c39438f5e3a731e3449f393634c8d5f0a043aae0eae327bf9350887ce8e6846a

Observation aa1b7741-f64a-4045-92a0-ab886997637b · outbound

This paper cites Clark, K.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Clark, K

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.194167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.125172Z digest=sha256:47ab6d49cbd2f3b4d6d58dac9cca3881606b4032bd0ba9f2537fc89be8e69073

Observation cad664e0-b2a2-46bc-b0f5-99639c46deb1 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.129784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.129784Z digest=sha256:7079f8946bf272909aff865d132641359aec01080c10720c1adac1c207bae37d

Observation 97370bb6-f807-4ab7-86d1-458051a8a23f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Training Verifiers to Solve Math Word Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.134983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.134983Z digest=sha256:20f2c6d1fa54c58832eb7500747bc12d778b950fe6261cadcbffe9a7b544158e

Observation 3e54834d-57eb-4a74-9042-3cf4dbc2279b · outbound

This paper cites Frohberg and F.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Frohberg and F

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.180081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.140039Z digest=sha256:9da9d640bc84e36b9b489ee17537bfba958c702d555e698cac57f5676b1d45b2

Observation 2b89a67e-665c-4c54-b7c8-78dc26687eee · outbound

This paper cites Guiver, S.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Guiver, S

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.166258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.144420Z digest=sha256:3a7d61ca38760b7ce6db56def126574455c8df6660362d6b10d053a1249e1ff9

Observation 1ae1f1f2-014a-4992-96a9-b8f3e2c3a9ab · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:03.152235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.149440Z digest=sha256:b232dcaefe1310ec179aaa656329ffcaf3abcdbdf522124fb74602a0e8be8046

Observation 094ceee2-e613-465d-9876-fff9a937eda6 · outbound

This paper cites Hendrycks, C.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Hendrycks, C

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.137847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.153996Z digest=sha256:80ee0609cac3b6d771f6a42ac5bb908b56b5ebb0ba19484ce495f5f5b59ace61

Observation 04a26570-1182-4b11-bdd9-7fe9326ba056 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.158497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.158497Z digest=sha256:e3764e6941f7dd1e22de3929be66b3a49a00e981f06162c53e8acdcc94e1a4fd

Observation 1701abe3-569f-4e95-ba65-170ea1c8c103 · outbound

This paper cites Kim et al.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Kim et al

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.124043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.163069Z digest=sha256:052c0365c04adeee6607d68da324214abdc04f076126305af840a479211b5a1b

Observation f47fa918-c496-460b-bb29-1aa8ff18da21 · outbound

This paper cites Kojima, S.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Kojima, S

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.110964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.167804Z digest=sha256:a216f0820be8fe9a1ef2b6110f85b2ff2303cf9bbdf311e82c895fb67750f18f

Observation d0e77fad-3de6-4f99-b2c2-a5e45e072299 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:03.096351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.172823Z digest=sha256:25c01bb82d1be30f701a5334b229fd0b5da2567f984a34b72cd591d5d87bd3af

Observation 25a96312-2b0b-475c-91b2-649f60b911e9 · outbound

This paper cites Evaluating the Robustness of Analogical Reasoning in Large Language Models.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Evaluating the Robustness of Analogical Reasoning in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.177268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.177268Z digest=sha256:82c2bea13bf93a564848b2a0986f86b2c05acf70dc08c0083e1d22c2d31d57bc

Observation 887ea66b-3234-48da-a262-04da90ae8d1f · outbound

This paper cites CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.182126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.182126Z digest=sha256:0fbc464b7b65d2f7eab28ced29b7508f75c429b62a36735004129518f9403fdf

Observation c0f0c23a-571e-4cbd-904c-acbbbaa255fd · outbound

This paper cites Lunardi, D.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Lunardi, D

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.081688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.186580Z digest=sha256:592536c0377daf3c854e76624446ceacb14af301d580cbd1e00bb9a7cce99391

Observation 96ccef58-4bca-4a9b-808f-ac8f72575dd8 · outbound

This paper cites Mihaylov, P.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Mihaylov, P

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.066683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.191311Z digest=sha256:8a2c0c59b51ea3a280d65bd582f7ce0de13d379f70a7efb98104086b37b54b26

Observation b1d0c527-9588-48b4-8bb9-4dbb8515e46e · outbound

This paper cites Mitchell.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Mitchell

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.052407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.196651Z digest=sha256:f0a7dd68575fb9eada5ed8f5ebb24d1c34a94d1dea31b73ba3d6b1fad68f5a40

Observation 226d59b7-7fb1-4c96-9e39-54c25d67787e · outbound

This paper cites Mitchell.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Mitchell

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.038500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.201503Z digest=sha256:53655b4ddbd205df80edca938c1582a71dd131bf07638395ce75b09fb560d8f5

Observation 33a6f282-cfff-4201-bc73-309166eb850f · outbound

This paper cites MS MARCO: A Human Generated MAchine Reading COmprehension Dataset.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs MS MARCO: A Human Generated MAchine Reading COmprehension Dataset

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.206347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.206347Z digest=sha256:264583a322f5c6df7f3e81959be4aa83cca54069d48680c0f8d7b6b5aff027ea

Observation fb30fad6-2087-4804-86dd-a88a403aa5df · outbound

This paper cites Ouyang, J.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Ouyang, J

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.024374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.211401Z digest=sha256:df02ecd315b51fbef4510a631041fee07c74dba2c1a7b6f2fc3fe2d29d1a93b5

Observation cd6581ea-42a0-4bde-8aa1-16a34cf5048f · outbound

This paper cites Variations in Relevance Judgments and the Shelf Life of Test Collections.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Variations in Relevance Judgments and the Shelf Life of Test Collections

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:31:02.633586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.215986Z digest=sha256:2bed7c84cca32f945830159809b02e74819f1d12247eb43b32e2d1186cf387fa

Observation 1928cce6-36f7-44eb-9b8a-6650c08abc75 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:03.010228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.220869Z digest=sha256:94173bee29b2599feb88bb5a3b4552228d7cdda17bd1a57a97c5257776b4346c

Observation 4a6e8c2e-a0c5-4393-8137-486c69e9eee8 · outbound

This paper cites Reuel-Lamparth, A.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Reuel-Lamparth, A

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.994446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.225854Z digest=sha256:f24a1c85caba743771f2920bdd3ff8d05948933e04699796733c043faf831539

Observation 9de30eba-ef39-4bd6-85ed-2ca5ebc5ff69 · outbound

This paper cites Sakaguchi, R.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Sakaguchi, R

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.979800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.230285Z digest=sha256:5580101f0298c5086b5c472867099046ec77261f81543e0a006aa96b8809dfc5

Observation 8644bea2-f5d8-45ff-8543-8e3c73b0cf41 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.234717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.234717Z digest=sha256:bb72fd40507eee579f9c087333811ce692d4838a34d34a5751187e95e0f5129a

Observation f473b916-536e-47ec-a003-aa68c32eebdf · outbound

This paper cites Sanderson.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Sanderson

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.964910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.238982Z digest=sha256:fd871462243f7299e5811585768697405e8ed5950723de34a5a07045915cb4f7

Observation c31df391-90e6-4fb8-a8b9-f824cca61ea2 · outbound

This paper cites Sclar, Y.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Sclar, Y

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.950379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.243182Z digest=sha256:0c73c360eaf6848069fc0941a3f18cb51ba17bcf56214f2145a923668e9395a6

Observation b87919ff-4415-45b9-9be5-d6d8977c26cc · outbound

This paper cites Sparck Jones and C.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Sparck Jones and C

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.935585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.248104Z digest=sha256:b20ee8a9cd4721addaaf207b719ae514531014665123351fef48a353c8d244c3

Observation b27103e1-f069-4ad6-83b9-9fbcef21a8a4 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.253189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.253189Z digest=sha256:ed43e965e351218280480aaf62864cbdf4c370a7c9037feb463c145505c8fa6a

Observation 49213a02-9fdd-4668-a7b4-068e583792e8 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:02.920333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.257873Z digest=sha256:683cbad03c3c9ec8307f5ff9f6b6292b1e154cabbf56610660b02ec75e5e58d3

Observation 3053029f-53cd-4dbf-b20d-ef2e25be86d0 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:02.903240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.262534Z digest=sha256:1245d52b8f77ca171f1ba7d38a5457c2332b642eb60262540999d97919c2e668

Observation 0ccf09db-f904-40b5-b722-508a892310dc · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:02.888353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.266985Z digest=sha256:99bb059304aab08041e0c38e3d0e02503d03f031c6c3b48fe37f9fa87e572c3c

Observation 1d473cba-0ea7-4dcb-88e8-b00755e160ae · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Finetuned Language Models Are Zero-Shot Learners

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.271888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.271888Z digest=sha256:269cb34c1ecdfabaac9ba89a17b9ed69f1d2a3e46e3842a8d7b4bcf849883215

Observation d2b7c9c3-2c39-49ea-bc90-c7c1ebba685c · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:02.873133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.276957Z digest=sha256:e50262b16d24ed4771fb40d5d957b3317493e32da0738e984bc4eb75443caa4b

Observation 5a247aea-7324-4277-b3d6-86a6d3e39e04 · outbound

This paper cites Emergent Abilities of Large Language Models.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Emergent Abilities of Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.282896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.282896Z digest=sha256:c108b870ba7c39394a58056438a6596fec695133c400daf5ff13d801e13031f3

Observation 826811b5-17ba-4d78-a2d8-1eebc9fcec3b · outbound

This paper cites Welbl, N.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Welbl, N

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.856887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.287920Z digest=sha256:29aff83de7cd68d3a6399924d334095689af5184b09fda8e808c4714b4f8c153

Observation bd8639d5-3661-49c9-9ae8-aeddd79afaae · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:02.842118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.292423Z digest=sha256:1fb210e49894aa3197b1fd8541406689a53645e3e2090fc89b3919d5898d4418

Observation e282fa3b-bb90-4bd8-b2d2-6c7cae47a607 · outbound

This paper cites Zellers, A.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Zellers, A

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.826115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.296921Z digest=sha256:a6d9b417de17dc9dad22e32a1d6cf5490349baabc1d949fd01c0f4adc6c7f7ab

Observation bc8a2d74-d951-4f55-ac91-916e6b2b04d7 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:02.811178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.301739Z digest=sha256:2c89db837d51c9026ec2a5a97cfbda020979e733c477df7a829d3a59536fce4f

Observation 4c487dba-c445-4e33-8390-0e3b677644ce · outbound

This paper cites Zheng et al.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Zheng et al

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.795648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:31:02.306286Z digest=sha256:d3faa9c64a010936317b71cd951782a1bd34477180f7de605266c441ecbef88d

Observation 01dc6a9e-3429-4a80-ae43-94fbd1f422be · outbound

This paper cites QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.310499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.310499Z digest=sha256:7c0714d71d943c849729a26551708a7339ebe5cc4400331f70318ef7f2ad300b

Observation 86795f54-c045-4a17-81a9-74e30fcc2ffc · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.314947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.314947Z digest=sha256:8a7fd734922f0cc94e3f97da979574b0e26a1953b1c9b48809cc8662e7c86926

Observation 61763fde-4f72-4bea-a209-5e3cf9f54dc0 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.319589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.319589Z digest=sha256:8d5b7bf757b1d9303f281345182a029767060f5f3f6bb280829c2537625665cc

Pith citing papers

Observation cd860ed7-1571-4add-9e0d-9db5135b3653 · inbound

Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users cites this paper.

Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:23:39.942350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T23:22:42.431997Z digest=sha256:33f5748d2e2a9110f0e595fe5713d3af10acc8628050cc55df6fcceeef8010c0

Observation a7bd56d7-f80a-4bfa-be9e-c15e06bc3b09 · inbound

StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario cites this paper.

StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:21:25.388199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T11:27:03.620795Z digest=sha256:384127301c8c3931d4f1cdd122c92782b4eb7909cfd41b73cd8a873b351dfb35

Observation b9447640-1d57-4f04-bb6a-4e4ce0214ec9 · inbound

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations cites this paper.

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:59.932740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T00:57:55.616036Z digest=sha256:6355183f47e426b754838bd8cbbddc6f9fe526b35d2266fd58855e5c86aca759

Observation 86bb03c9-9819-40e2-95cb-6390ff59ae8e · inbound

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations cites this paper.

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.000489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-30T23:42:00.965554Z digest=sha256:867e9606d488574a5eb079d1c4279d6d0b1a19d547c7d0b2edcea110d6c6fa1f

Observation 8e54f074-d4a3-4308-8237-8717c9608525 · inbound

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning cites this paper.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.325279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:7afb4b7111aabd396fe4e0194bd1bb4783935f02a28d49c4b8ff388e2f7d3963

Observation b0a32282-5a01-4201-ade7-b59c79931a23 · inbound

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments cites this paper.

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:53:40.568253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T16:51:36.524194Z digest=sha256:e7805394dc366cf61d2f9e6ee0452983f48eeb45c0891b834ecbf80c26124d57

Observation cacb1980-62da-4e63-b519-bb6e4d88023a · inbound

MAVEN: Improving Generalization in Agentic Tool Calling cites this paper.

MAVEN: Improving Generalization in Agentic Tool Calling On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.135066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:437129f74d9f8495a6e61312dd2600cc24752dadf2032f334467098e6a6e96ad

Observation b9caf502-53cd-465a-a0b2-1d1934c03339 · inbound

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy cites this paper.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:08.844007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:08.844007Z digest=sha256:904251f2cc5307d4e8568affccdbd8f56fc72e86dc2b7a3700ee69de39253637