Pith. sign in

Paper Citation Record · LEDGER

Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2406.07545.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.07545 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:41:40.029820Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:17:57.221151Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cbee2a8d-6a73-4c03-bd79-2bfedb1e788f · inbound

VoiceBench: Benchmarking LLM-Based Voice Assistants cites this paper.

VoiceBench: Benchmarking LLM-Based Voice Assistants Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:50:14.000184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T00:50:13.841689Z digest=sha256:36dd7d6a0079e995acdd17a29ea3b82eaf6790ec84f31a2cc3f670f596ccac1f

Observation c79cca48-3ef8-47d3-b589-8fb08bb09684 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:13.386774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:78af9fc6527901982a922f8f1c572d6510aa07985319cf143da5d678d1caaf12

Observation d031f4f1-005b-4baf-8be1-021ff3e3a4fa · inbound

Human-Centric Evaluation for Foundation Models cites this paper.

Human-Centric Evaluation for Foundation Models Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.030772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.030772Z digest=sha256:57f89ba53132b79b7acf4de4564ac17e447eb1d75d4f395d058d517f01123760

Observation 47cbe7f8-93d0-48a8-82ed-ba153745657c · inbound

DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation cites this paper.

DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:57.059533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:57.059533Z digest=sha256:39898dfaa4073cf56b6f033e45aa61adce71b55a3363d13d063ecb36ea88cce8

Observation 2a45c7a3-b8cd-497e-8811-e2f50673d835 · inbound

Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans cites this paper.

Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:40.029820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:41:40.029820Z digest=sha256:87a4f0c70e85003d092c394d9e48dd23e4a2e50e0537eeba48c402989f50f348

Observation 8fe4b995-1c8b-407f-b102-12b70a3744c7 · inbound

SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version cites this paper.

SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T17:57:17.759571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:57:17.759571Z digest=sha256:344931e903bc188f839a41d3bd54e304810d678d9258035047f7a24183bbf6c2

Observation 7727e9bb-03d0-4a6f-908e-27a44bdca87b · inbound

Position: AI Evaluations Should be Grounded on a Theory of Capability cites this paper.

Position: AI Evaluations Should be Grounded on a Theory of Capability Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:00:41.530217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T21:57:55.834632Z digest=sha256:b725c84849166e9c34021d6eac0a18deb67db836c4daf8b352efff83ddc5d7af

Observation 437054b5-47cb-4991-b811-c6b8ac66e482 · inbound

Efficient Evaluation of LLM Performance with Statistical Guarantees cites this paper.

Efficient Evaluation of LLM Performance with Statistical Guarantees Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:00:51.716995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T10:58:40.958435Z digest=sha256:88c656abec88249bdb1897f4f886fd2dce4acdfd08a47a5d20b72f722344c50f

Observation 9abee72f-ed8b-4cc7-bed1-4b05ec0a9710 · inbound

KNIGHT: Knowledge Graph-Driven Multiple-Choice Question Generation with Adaptive Hardness Calibration cites this paper.

KNIGHT: Knowledge Graph-Driven Multiple-Choice Question Generation with Adaptive Hardness Calibration Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:05.759485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:05.759485Z digest=sha256:0b3a2e83501cc1ba0fe03643d2f6be256d3e8bae64fc347a3281a2e41bde56a8

Observation b9726ac4-c6f7-42ef-9a95-0595c74e9aae · inbound

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety cites this paper.

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-15T13:17:48.274611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:17:48.274611Z digest=sha256:5d194fa6cb3bf0d80c11ff1fb3951da028c73796bfbc06772d6c553c1f640c4e

Observation 5a9037bc-3af2-4654-baba-8e4918c7f508 · inbound

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models cites this paper.

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:08:25.171011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T23:07:38.352198Z digest=sha256:a0b937ae38aacc3e1095882cb5339deb14d99c5d9e5f7408071d639b34e0ddbd

Observation 9cd40c27-7468-4b36-8464-767b41362652 · inbound

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models cites this paper.

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T17:05:10.686785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T17:05:10.686785Z digest=sha256:c8f547a8c888ce77701604717723f7b5d2447598aea624612e582da940f01906

Observation 7d54dfb1-535a-410c-bd23-840e74d9d914 · inbound

Improving Cross-Format Robustness in Language Models with Multi-Format Training cites this paper.

Improving Cross-Format Robustness in Language Models with Multi-Format Training Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:17:57.222547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:13:22.927828Z digest=sha256:ab323a45ee0925befb21785db2d3cdc7ebbad6cb8c0f4d8e5835194d17a31f1e