Pith. sign in

Paper Citation Record · LEDGER

Statistical multi-metric evaluation and visualization of LLM system predictive performance

As of 19 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2501.18243.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18243 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:16:37.811235Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T09:45:26.654826Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:47:27.686727Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 566efa22-b0e5-4cca-b077-c4d02f84bf7f · outbound

This paper cites Using Combinatorial Optimization to Design a High quality LLM Solution.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Using Combinatorial Optimization to Design a High quality LLM Solution

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-10T00:16:37.855138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.732561Z digest=sha256:12e94cf00d5c52742472ee7d9f7f578d29f44e1d4557f6b3630e444539eb4f55

Observation 04234310-eb59-4515-bb0d-50ff4bb4e5e1 · outbound

This paper cites Simultaneous confidence intervals for ranks with application to ranking institutions.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Simultaneous confidence intervals for ranks with application to ranking institutions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.068354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.736883Z digest=sha256:f5382a95bb837f202d2f6715169980bae3293e40f1eac23e260a1e3b6ae16bdf

Observation dbba6ae6-80c4-4a9a-aefb-711f4223e140 · outbound

This paper cites Statistical Power Analysis for the Behavioral Sciences.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Statistical Power Analysis for the Behavioral Sciences

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.057584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.740522Z digest=sha256:65ef42c6d7e8c334197ac0873bf2e472a1bbeb689ce8499a11e9b04237c336da

Observation 65ecd948-64de-421c-b42e-bc66b4a359c7 · outbound

This paper cites The spotis rank reversal free method for multi-criteria decision-making support.

Statistical multi-metric evaluation and visualization of LLM system predictive performance The spotis rank reversal free method for multi-criteria decision-making support

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.047600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.743949Z digest=sha256:47bf5ccff9f87eac08b182a08f9a0b7befcbb158c181ded01677909c0726cc51

Observation 146b089a-d610-4ffb-9ddd-3d33d9cc0478 · outbound

This paper cites CrossCodeEval : A diverse and multilingual benchmark for cross-file code completion.

Statistical multi-metric evaluation and visualization of LLM system predictive performance CrossCodeEval : A diverse and multilingual benchmark for cross-file code completion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.037414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.747876Z digest=sha256:00e30397f572aa5793eeb59b4cb5d8514028fc79c0e9119e4e71f8788d8fb308

Observation 10d7a8cd-a604-4d69-8e87-deff4db695cc · outbound

This paper cites Multiobjective optimization in river basin development.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Multiobjective optimization in river basin development

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.027060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.751357Z digest=sha256:5777dbd15a7853b26afd999abbb73507faf087394c610f642a9905c4dfecc5ee

Observation 18e66ac1-1997-42f8-af99-510f415f4646 · outbound

This paper cites Sensitivity of decisions to probability estimation errors: A reexamination.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Sensitivity of decisions to probability estimation errors: A reexamination

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.016971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.754939Z digest=sha256:6f307ce774f3573e6459438d5179a4498be76fcb57557d634c99b5e0049c8a2e

Observation a3cb713e-bac8-4366-ba79-3b15c89ed2ac · outbound

This paper cites Performances are plateauing, let's make the leaderboard steep again, 2024.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Performances are plateauing, let's make the leaderboard steep again, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.006862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.758183Z digest=sha256:2e7393dc186878d9920990bee7b59307058009e2dbb5897ddaf86f06f467edeb

Observation 3924f555-4d03-4f13-a201-3e123cb7d0b7 · outbound

This paper cites Choosing between methods of combining-values.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Choosing between methods of combining-values

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.996655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.761668Z digest=sha256:e39d5b7060c03fd36ffb5fbedd4573b9a1958e30d47d60dcc0f6c16f4d458aaa

Observation 6ead93d0-6238-4d15-90c9-47eb85b512fb · outbound

This paper cites Methods for Multiple Attribute Decision Making, pages 58--191.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Methods for Multiple Attribute Decision Making, pages 58--191

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T00:16:37.765034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:16:37.765034Z digest=sha256:cdd1092e6c04dcccdb339cf5988fa0fdfad06db8d9d8e63ac564e8fb45621a85

Observation 7d9e3fec-4b19-4467-902d-03cbffc60de0 · outbound

This paper cites pymcdm—the universal library for solving multi-criteria decision-making problems.

Statistical multi-metric evaluation and visualization of LLM system predictive performance pymcdm—the universal library for solving multi-criteria decision-making problems

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.986406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.768698Z digest=sha256:e3683bb3d0975add2b96ce40b5bf07c452154bcb0032671fb733e4c35c6f5d6e

Observation fb46097e-e7f7-4c1c-833a-96b88c2b8ee0 · outbound

This paper cites an unresolved cited work.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Unresolved cited work

Reference 12

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T00:16:37.976123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.771984Z digest=sha256:ec1ded27af3141ee469e8b16b0f54e188556d9716f27c0360d775de9a68d439c

Observation 80e6f8df-2f84-4966-902d-3c5a65cb3b2e · outbound

This paper cites an unresolved cited work.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:16:37.965615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.775570Z digest=sha256:998040c59ee8ff81713198b6ef7f8a2194089012a4a766989835a5f7c7c108e6

Observation e0dde2ef-6946-42d1-a5ff-4b4f2838527c · outbound

This paper cites Mangiafico.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Mangiafico

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.955295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.778900Z digest=sha256:ae5897a895ad4a36bf7e3648714d49433634b307cd15e1f5337903c186a51850

Observation fea60797-81dd-4a12-b480-25545207c333 · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Granite Code Models: A Family of Open Foundation Models for Code Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T00:16:37.782195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:16:37.782195Z digest=sha256:382a31147c5d395bc9e2d4bd1c159c0db17e262ce55b7625974ad8d22d7c3298

Observation 6cd96f46-fd17-454c-851c-7f8ef76e2f37 · outbound

This paper cites New effect size rules of thumb.

Statistical multi-metric evaluation and visualization of LLM system predictive performance New effect size rules of thumb

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.944589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.785701Z digest=sha256:b28bdd22ada2cd1265b56c368e5cdcb83b47391ea2ac8ff0720e4cd8539442dc

Observation 38214852-9c4a-4136-8253-42363027f9b3 · outbound

This paper cites Statsmodels: Econometric and statistical modeling with python.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Statsmodels: Econometric and statistical modeling with python

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.933028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.788961Z digest=sha256:71da9af9acea9b120cd876b900f82ce5cb1bfd3599fe63a961d312554184a761

Observation 10c01e60-8dca-4ec2-80f1-0f73ea3f7893 · outbound

This paper cites A multiple criteria decision making method based on relative value distances.

Statistical multi-metric evaluation and visualization of LLM system predictive performance A multiple criteria decision making method based on relative value distances

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.923503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.792056Z digest=sha256:7f82d45c0875abdd11c0ebd6519b246c8bb9cac20816f5a471104e44d663ffd3

Observation e09f9612-cc96-4cb7-97eb-66ac0de49479 · outbound

This paper cites Comparative analysis of some prominent mcdm methods: A case of ranking serbian banks.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Comparative analysis of some prominent mcdm methods: A case of ranking serbian banks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.914021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.795146Z digest=sha256:852cf35e604459e46e872bb9f434c2b6a52e3ce1adac0449b41636ed2eee9627

Observation 2befde40-c5ff-4ec2-b8a0-d120decf3582 · outbound

This paper cites Using effect size—or why the p value is not enough.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Using effect size—or why the p value is not enough

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.904452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.798176Z digest=sha256:e736cfbd145ae91178e41dfba1d828015e1dd827a31c00a14d34b2323ca89fb5

Observation cc56db27-137e-4e46-8edc-bc2589e04b89 · outbound

This paper cites Calculating and synthesizing effect sizes.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Calculating and synthesizing effect sizes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.894808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.802219Z digest=sha256:1d68f18b47c317966fbdeabae37b37a9c9cd43f371277123cdabee2c2011f7bd

Observation ebff8b2c-07e0-4a4a-acb6-6d151fa38c95 · outbound

This paper cites Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, St \'e fan J.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, St \'e fan J

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T00:16:37.805163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:16:37.805163Z digest=sha256:35a35ff0f00366c01ec2c59f19f858e3c35dcbdbbb063b41e28355fa8b544706

Observation 616e18c3-5041-4049-9a04-c05c459df20e · outbound

This paper cites The harmonic mean p-value for combining dependent tests.

Statistical multi-metric evaluation and visualization of LLM system predictive performance The harmonic mean p-value for combining dependent tests

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.876952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.808144Z digest=sha256:ad82252afd00e10d2deee158209cf656ad40de5c254bf16d6a5bf69be8be42b8

Observation 651f04b1-56d5-49d7-bca7-d10048e76774 · outbound

This paper cites harmonicmeanp tutorial, 2024.

Statistical multi-metric evaluation and visualization of LLM system predictive performance harmonicmeanp tutorial, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.866262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.811235Z digest=sha256:e20c7695c0ad556959a5a7d67499e82a899227cf874d200af216a1a5b7a3ca54

Pith citing papers

Observation 5e68a197-fb9c-432f-9662-1cf1e06d0e88 · inbound

Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation cites this paper.

Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation Statistical multi-metric evaluation and visualization of LLM system predictive performance

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:47:27.688270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T17:54:22.974336Z digest=sha256:d87e074dea0783adbd4bfd0f1b7d74c0b0ee50c06bb6f3705427f4df7cef9b14

Observation 0ce2d397-8c5e-43ab-b500-c7fbe4be67dd · inbound

Quantifying Ranking Uncertainty in LLM Benchmarks cites this paper.

Quantifying Ranking Uncertainty in LLM Benchmarks Statistical multi-metric evaluation and visualization of LLM system predictive performance

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T09:45:26.654826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:45:26.654826Z digest=sha256:383622773669cb95ee1f6994a560ea8f7c5d33947b002f326e69c1bf0602471e