Pith. sign in

Paper Citation Record · LEDGER

Statistical multi-metric evaluation and visualization of LLM system predictive performance

As of 11 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2501.18243.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18243 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:16:37.811235Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T09:45:26.654826Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:47:27.686727Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 566efa22-b0e5-4cca-b077-c4d02f84bf7f · outbound

This paper cites Using Combinatorial Optimization to Design a High quality LLM Solution.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Using Combinatorial Optimization to Design a High quality LLM Solution

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-10T00:16:37.855138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.732561Z digest=sha256:2eebaa714097ac618689ec9a65f2be438e8161e0c5eac87ae4f783a57c5a1b8c

Observation 04234310-eb59-4515-bb0d-50ff4bb4e5e1 · outbound

This paper cites Simultaneous confidence intervals for ranks with application to ranking institutions.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Simultaneous confidence intervals for ranks with application to ranking institutions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.068354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.736883Z digest=sha256:794cc6c0e09c191f5092cbbb0ef0699826b2dda326231d55759dd6c63504de9e

Observation dbba6ae6-80c4-4a9a-aefb-711f4223e140 · outbound

This paper cites Statistical Power Analysis for the Behavioral Sciences.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Statistical Power Analysis for the Behavioral Sciences

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.057584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.740522Z digest=sha256:13783c684da64486067ccf3ca242ceb25e04a9d0237e597764b67a1fac25fcca

Observation 65ecd948-64de-421c-b42e-bc66b4a359c7 · outbound

This paper cites The spotis rank reversal free method for multi-criteria decision-making support.

Statistical multi-metric evaluation and visualization of LLM system predictive performance The spotis rank reversal free method for multi-criteria decision-making support

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.047600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.743949Z digest=sha256:2445050647f743842cf33f5235d4611609d5ba4be1d98a276ae3d924b339fa57

Observation 146b089a-d610-4ffb-9ddd-3d33d9cc0478 · outbound

This paper cites CrossCodeEval : A diverse and multilingual benchmark for cross-file code completion.

Statistical multi-metric evaluation and visualization of LLM system predictive performance CrossCodeEval : A diverse and multilingual benchmark for cross-file code completion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.037414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.747876Z digest=sha256:fe3962df687ba09666479f10fea68be593638aaf52200eda61866c599ca66282

Observation 10d7a8cd-a604-4d69-8e87-deff4db695cc · outbound

This paper cites Multiobjective optimization in river basin development.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Multiobjective optimization in river basin development

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.027060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.751357Z digest=sha256:0876a29b34af89226cd2f2cbe187aaf69a7509fa1b110f76dadbe99a84164453

Observation 18e66ac1-1997-42f8-af99-510f415f4646 · outbound

This paper cites Sensitivity of decisions to probability estimation errors: A reexamination.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Sensitivity of decisions to probability estimation errors: A reexamination

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.016971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.754939Z digest=sha256:7e534449c59bb134e2b5081a103a4295627c26b6ecdc2444cc396c762d449f83

Observation a3cb713e-bac8-4366-ba79-3b15c89ed2ac · outbound

This paper cites Performances are plateauing, let's make the leaderboard steep again, 2024.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Performances are plateauing, let's make the leaderboard steep again, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.006862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.758183Z digest=sha256:6910e23076772c4a8160603d207a637fbfae2125bcb84f235b48ea240246cc82

Observation 3924f555-4d03-4f13-a201-3e123cb7d0b7 · outbound

This paper cites Choosing between methods of combining-values.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Choosing between methods of combining-values

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.996655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.761668Z digest=sha256:ae694522edb50d61e72722d4e1653231f6f962b844ec4c82a497bae478db5f6e

Observation 6ead93d0-6238-4d15-90c9-47eb85b512fb · outbound

This paper cites Methods for Multiple Attribute Decision Making, pages 58--191.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Methods for Multiple Attribute Decision Making, pages 58--191

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T00:16:37.765034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:16:37.765034Z digest=sha256:f14f5a061a02a92a8c14cc4a1431940776b7a9f03fa9aa85c80982f714958be7

Observation 7d9e3fec-4b19-4467-902d-03cbffc60de0 · outbound

This paper cites pymcdm—the universal library for solving multi-criteria decision-making problems.

Statistical multi-metric evaluation and visualization of LLM system predictive performance pymcdm—the universal library for solving multi-criteria decision-making problems

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.986406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.768698Z digest=sha256:ba6fc71396924df2b1d08eb5e3553652c6e12f395ad54b6ca180e3cee4873e22

Observation fb46097e-e7f7-4c1c-833a-96b88c2b8ee0 · outbound

This paper cites an unresolved cited work.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Unresolved cited work

Reference 12

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T00:16:37.976123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.771984Z digest=sha256:add3bcfcff6b917178268f94e0c6ebe134a3994969b0c4e942662ef88e3c895e

Observation 80e6f8df-2f84-4966-902d-3c5a65cb3b2e · outbound

This paper cites an unresolved cited work.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:16:37.965615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.775570Z digest=sha256:5fcd7484ca1a8bcb90fa1cc9a672a29de478c386873798582855eac6bc20399b

Observation e0dde2ef-6946-42d1-a5ff-4b4f2838527c · outbound

This paper cites Mangiafico.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Mangiafico

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.955295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.778900Z digest=sha256:284ed02d33299813e895bc4a5b9bf545b115840b9102cf729ea87af6e18ee772

Observation fea60797-81dd-4a12-b480-25545207c333 · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Granite Code Models: A Family of Open Foundation Models for Code Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T00:16:37.782195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:16:37.782195Z digest=sha256:41097ab88afbeef17829becc3e40f54d774b43f9feb937814c4f507b14e26635

Observation 6cd96f46-fd17-454c-851c-7f8ef76e2f37 · outbound

This paper cites New effect size rules of thumb.

Statistical multi-metric evaluation and visualization of LLM system predictive performance New effect size rules of thumb

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.944589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.785701Z digest=sha256:581a21cf84cf868258940e67eb4dc0f1da60cd9844f40551c40d256dbfb1d4ab

Observation 38214852-9c4a-4136-8253-42363027f9b3 · outbound

This paper cites Statsmodels: Econometric and statistical modeling with python.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Statsmodels: Econometric and statistical modeling with python

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.933028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.788961Z digest=sha256:41359b7fa10188fff434c9e4e1f16181fcb4427d33ae026ecb345cfb00c4d372

Observation 10c01e60-8dca-4ec2-80f1-0f73ea3f7893 · outbound

This paper cites A multiple criteria decision making method based on relative value distances.

Statistical multi-metric evaluation and visualization of LLM system predictive performance A multiple criteria decision making method based on relative value distances

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.923503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.792056Z digest=sha256:d647b171a6d15ed6f2374427c76700595e5e91c12cad8027d7e25c7170f3934c

Observation e09f9612-cc96-4cb7-97eb-66ac0de49479 · outbound

This paper cites Comparative analysis of some prominent mcdm methods: A case of ranking serbian banks.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Comparative analysis of some prominent mcdm methods: A case of ranking serbian banks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.914021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.795146Z digest=sha256:bcb9ee50bd6e2eb93e521a60a5a8828c0455bf7a95df9522ba01d8b1f4097ad7

Observation 2befde40-c5ff-4ec2-b8a0-d120decf3582 · outbound

This paper cites Using effect size—or why the p value is not enough.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Using effect size—or why the p value is not enough

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.904452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.798176Z digest=sha256:f926e558bc932fdae42294acbe6834e56b1df1a17b5d8d48fb5fbd0d7430d934

Observation cc56db27-137e-4e46-8edc-bc2589e04b89 · outbound

This paper cites Calculating and synthesizing effect sizes.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Calculating and synthesizing effect sizes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.894808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.802219Z digest=sha256:63ab8a6585d2440351ee08658574246d056d3093ababee3de47da3939b5fcb23

Observation ebff8b2c-07e0-4a4a-acb6-6d151fa38c95 · outbound

This paper cites Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, St \'e fan J.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, St \'e fan J

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T00:16:37.805163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:16:37.805163Z digest=sha256:28170bff0d2fc54903f8de58156c384b2238849f1f459fd59be0bb7e4b26a9df

Observation 616e18c3-5041-4049-9a04-c05c459df20e · outbound

This paper cites The harmonic mean p-value for combining dependent tests.

Statistical multi-metric evaluation and visualization of LLM system predictive performance The harmonic mean p-value for combining dependent tests

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.876952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.808144Z digest=sha256:c9f3030bd9d92df5785bb285aa85c59f0d0fb4b12fd3eb04880de9403cec8b62

Observation 651f04b1-56d5-49d7-bca7-d10048e76774 · outbound

This paper cites harmonicmeanp tutorial, 2024.

Statistical multi-metric evaluation and visualization of LLM system predictive performance harmonicmeanp tutorial, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.866262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.811235Z digest=sha256:d70df37b6f1e0b759f0428a64bf88eb3d7f1eda500c7dd162d6fc47eae4c3f88

Pith citing papers

Observation 5e68a197-fb9c-432f-9662-1cf1e06d0e88 · inbound

Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation cites this paper.

Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation Statistical multi-metric evaluation and visualization of LLM system predictive performance

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:47:27.688270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T17:54:22.974336Z digest=sha256:5576f7941a0d556ca2ca29c4d5d700b5611df6cfafb16d61f5024e79787b37dc

Observation 0ce2d397-8c5e-43ab-b500-c7fbe4be67dd · inbound

Quantifying Ranking Uncertainty in LLM Benchmarks cites this paper.

Quantifying Ranking Uncertainty in LLM Benchmarks Statistical multi-metric evaluation and visualization of LLM system predictive performance

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T09:45:26.654826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:45:26.654826Z digest=sha256:e63a2af5ebc4eb7e7f8a8899d3e4f70cda04215c8a773dd51c37692435680174