Pith. sign in

Paper Citation Record · LEDGER

Statistical Multicriteria Evaluation of LLM-Generated Text

As of 17 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 1 inbound Pith citation observation for arXiv:2506.18082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18082 v2

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:01:23.023749Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T04:49:15.239636Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T04:52:17.135964Z

Reference resolution

86 of 86 outbound references displayed

  • verified exact5
  • verified fuzzy39
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 73934b5a-b8bd-4fad-9184-f14341d7d33b · outbound

This paper cites GPT-4 Technical Report.

Statistical Multicriteria Evaluation of LLM-Generated Text GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.612704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.612704Z digest=sha256:7c5ad5b9ccd7e851c5c0fe84c6f3a12e3943b8b9fed95c0ea19fa214fbbae302

Observation 047cc169-9748-4ac8-8dd0-afd73b65e784 · outbound

This paper cites A learning algorithm for boltzmann machines.

Statistical Multicriteria Evaluation of LLM-Generated Text A learning algorithm for boltzmann machines

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.618735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.618735Z digest=sha256:59e1c010539babc399ee141edc76e4f7241246d38f90e48213eb0343dc004c1b

Observation 3d5a76c7-339b-4488-b6d2-3deb50423093 · outbound

This paper cites Bayesian Optimization for Building Social-Influence-Free Consensus.

Statistical Multicriteria Evaluation of LLM-Generated Text Bayesian Optimization for Building Social-Influence-Free Consensus

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.623798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.623798Z digest=sha256:5bf71d273bba6c0d69e29c983836e1717130613b6c8e47b7c07c69bff0dadc0c

Observation bd10c74c-bbdc-4b18-9328-500dd98a5989 · outbound

This paper cites Jointly measuring diversity and quality in text generation models.

Statistical Multicriteria Evaluation of LLM-Generated Text Jointly measuring diversity and quality in text generation models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.629098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.629098Z digest=sha256:d74bd8630a26d8687df4be94143a3111857e35225aecdd18a29d2035bd5c52cb

Observation 6c93fbea-aa3e-477e-b6c0-6075fae97fbc · outbound

This paper cites Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges.

Statistical Multicriteria Evaluation of LLM-Generated Text Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.634347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.634347Z digest=sha256:3c7a18415285c0bcb49a95e79b04638f9b48602f0936f32879807c4306e4e5dc

Observation 6ab273dd-3369-4e8b-8c2c-519065bcb098 · outbound

This paper cites Howcroft.

Statistical Multicriteria Evaluation of LLM-Generated Text Howcroft

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.639624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.639624Z digest=sha256:57f8e85adbe79cd8073ec84bab8a5ed5cb5472faf770099ff132ccae819bf346

Observation d80f3a6b-10a5-4745-abef-2f56854447b3 · outbound

This paper cites Time for a change: a tutorial for comparing multiple classifiers through bayesian analysis.

Statistical Multicriteria Evaluation of LLM-Generated Text Time for a change: a tutorial for comparing multiple classifiers through bayesian analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.645123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.645123Z digest=sha256:f04b31276b51c53f384583f1a59fb7bc9c47cccf32cf609596094df942a1c4fb

Observation be63f9a0-102e-48ad-ad99-ea2ee7f1b679 · outbound

This paper cites Estimating the replication probability of significant classification benchmark experiments.

Statistical Multicriteria Evaluation of LLM-Generated Text Estimating the replication probability of significant classification benchmark experiments

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.649563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.649563Z digest=sha256:4dc0b5db350d1cb18fb7ac06a2e7ab8daf976eb1166c8e7dd5187bcc9417ef88

Observation 180392b0-bcbb-4614-b21d-5cc702d4a68f · outbound

This paper cites Comparing machine learning algorithms by union-free generic depth.

Statistical Multicriteria Evaluation of LLM-Generated Text Comparing machine learning algorithms by union-free generic depth

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.654122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.654122Z digest=sha256:0c056da61633951c2672154c52c179221a6fb9fe77b09edacbae719ddd8d8854

Observation 7fd33ca0-421c-428a-b184-435ac2a8c915 · outbound

This paper cites How to Choose a Reinforcement-Learning Algorithm.

Statistical Multicriteria Evaluation of LLM-Generated Text How to Choose a Reinforcement-Learning Algorithm

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:01:23.498287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.659145Z digest=sha256:7e744e132fd651abea50ab282de13f79f5f9b898e9dd9788d4d193feec3b5a6e

Observation 2f2a6f9b-d7d4-4e06-85b5-1e658d827352 · outbound

This paper cites Self-learning from pairwise credal labels.

Statistical Multicriteria Evaluation of LLM-Generated Text Self-learning from pairwise credal labels

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.664456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.664456Z digest=sha256:5c4b5f897f5759311dcf85e1cce1c43fdf6332ca576a8440b21129d2016eecb4

Observation 06115b3f-35d0-41a9-92d9-a41461efc11c · outbound

This paper cites Credal Bayesian Deep Learning.

Statistical Multicriteria Evaluation of LLM-Generated Text Credal Bayesian Deep Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.668977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.668977Z digest=sha256:77d62c55746d090c090083e499853a28bcb2cb526a48e21baa2b7368e28a2e17

Observation a7374a2e-e032-4ad9-b368-1b027ee2837b · outbound

This paper cites Evaluation of Text Generation: A Survey.

Statistical Multicriteria Evaluation of LLM-Generated Text Evaluation of Text Generation: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.674048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.674048Z digest=sha256:2f50f23cb41233b2bc8326d52932c9fa9f24d3f5ee28b061ca166089e4e4b9d9

Observation 37a75dca-7e3d-4efb-968e-8ccfe098d594 · outbound

This paper cites Evaluating language models as risk scores.

Statistical Multicriteria Evaluation of LLM-Generated Text Evaluating language models as risk scores

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.679008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.679008Z digest=sha256:900faa130383a346b407e64de3732f31b1df010878f66e05533c4bc237382f21

Observation de2d369e-ba1e-41c7-83f6-96f86b9678c6 · outbound

This paper cites Dem s ar.

Statistical Multicriteria Evaluation of LLM-Generated Text Dem s ar

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.684349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.684349Z digest=sha256:e4a073636da484669c3efb697b750d794825a6aa38bea9780bb6ac6326607678

Observation 48152cc6-be62-424b-90c2-7df5bd442a12 · outbound

This paper cites Semi-supervised learning guided by the generalized bayes rule under soft revision.

Statistical Multicriteria Evaluation of LLM-Generated Text Semi-supervised learning guided by the generalized bayes rule under soft revision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.370372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.688853Z digest=sha256:8a5b1b28b80b8a3f0bc19418a6f85f838f1fcb133fae9f93d8c357096944bbef

Observation 50948c9c-2caf-48cc-89a4-cd79f0608d6f · outbound

This paper cites an unresolved cited work.

Statistical Multicriteria Evaluation of LLM-Generated Text Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:01:24.355254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.693271Z digest=sha256:3d40d8b27737b94b4fa09bcfb32ce199c7326254c95df094b0ac2882aedff415

Observation 5d176290-587d-4e28-8a29-41249b8a12a9 · outbound

This paper cites Eugster, T.

Statistical Multicriteria Evaluation of LLM-Generated Text Eugster, T

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.340072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.699076Z digest=sha256:6fa6d64fb62b9f2712b41e4535c399df65162dfd6dd4187c6759a483bed73670

Observation 0b2b8622-edac-4319-8f46-9928f3308b49 · outbound

This paper cites Hierarchical neural story generation, 2018.

Statistical Multicriteria Evaluation of LLM-Generated Text Hierarchical neural story generation, 2018

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.323904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.703565Z digest=sha256:e01089ea7680a6bb2adde6f0e497539725dc39da141ade48403c5caf8c0fcfc5

Observation a470a5ca-8b39-47e1-9cfa-ce2230e1f604 · outbound

This paper cites Beam search strategies for neural machine translation.

Statistical Multicriteria Evaluation of LLM-Generated Text Beam search strategies for neural machine translation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.708202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.708202Z digest=sha256:70f265f8339b14361951c8da9c53160ec7004d983c36411fd4adae2f6769f3c2

Observation 502d7b1d-5c40-4a6f-8b77-93039ee13f03 · outbound

This paper cites Simcse: Simple contrastive learning of sentence embeddings, 2022.

Statistical Multicriteria Evaluation of LLM-Generated Text Simcse: Simple contrastive learning of sentence embeddings, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.712984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.712984Z digest=sha256:e9d85ac2aced9e494d293dbf0936b70c410221cc78372a3b1a23dd1c94386601

Observation c70d8371-79a8-4660-aed6-08e2de291b8b · outbound

This paper cites Adaptive Contrastive Search: Uncertainty-Guided Decoding for Open-Ended Text Generation.

Statistical Multicriteria Evaluation of LLM-Generated Text Adaptive Contrastive Search: Uncertainty-Guided Decoding for Open-Ended Text Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.717493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.717493Z digest=sha256:0333bce2ca369ecd5f265b2721d3ae94a0ae9daa707ca14787fa3f8942466491

Observation 809d74bb-aed3-46c0-b02c-1d76cb139846 · outbound

This paper cites Decoding decoded: Understanding hyperparameter effects in open-ended text generation.

Statistical Multicriteria Evaluation of LLM-Generated Text Decoding decoded: Understanding hyperparameter effects in open-ended text generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.295768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.722044Z digest=sha256:b536ffd295da565d3939665c9aeb84803e098867978ac7016681bcf2e76e558d

Observation 41b2dbed-66f3-4aca-8dca-3aeb8f0d9da1 · outbound

This paper cites Towards Better Open-Ended Text Generation: A Multicriteria Evaluation Framework.

Statistical Multicriteria Evaluation of LLM-Generated Text Towards Better Open-Ended Text Generation: A Multicriteria Evaluation Framework

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:01:23.406919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.726512Z digest=sha256:e41656f808b42dc06f6dd541fce172ef4fda30dd78832181c2cb62b82e111f30

Observation 47b56d44-0696-4371-ae7c-de8d090e9441 · outbound

This paper cites Garc \'i a and F.

Statistical Multicriteria Evaluation of LLM-Generated Text Garc \'i a and F

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.280452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.731330Z digest=sha256:6fa33912b3b7f7dbcc4c15475c95e3bc352fe461504f1d5f16bc1c95f1abf5c2

Observation 98af811d-5f3d-42e5-8310-aefcfe993338 · outbound

This paper cites García, A.

Statistical Multicriteria Evaluation of LLM-Generated Text García, A

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.266075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.736083Z digest=sha256:6addb2e04888a6c07d538bf73f08b87a6f06c2a60eb29bc04940cf4c250296cb

Observation 847f3e74-df78-4e72-835d-c74db774bf9d · outbound

This paper cites The Llama 3 Herd of Models.

Statistical Multicriteria Evaluation of LLM-Generated Text The Llama 3 Herd of Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.740637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.740637Z digest=sha256:1729ff5a7a24c1de1918593c0c59ec2862343aa0d106ac0ba6c1adda8c0864c4

Observation 62970aa5-b0fd-448e-87e5-f8fdbb56fd96 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Statistical Multicriteria Evaluation of LLM-Generated Text DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.745434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.745434Z digest=sha256:5ae87af8e12329f51afec3ddeaeda167993203d71ae452c9eba1d9b682b4c6c7

Observation f06a955f-23d7-49ff-9010-f1e8e2bd8cf0 · outbound

This paper cites Unifying Human and Statistical Evaluation for Natural Language Generation.

Statistical Multicriteria Evaluation of LLM-Generated Text Unifying Human and Statistical Evaluation for Natural Language Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.749990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.749990Z digest=sha256:dc5fdb9ba8f791cc808a9270b80dc08507ebb8ff7ae08f6ab1173552042d5a83

Observation d9a87aa2-064c-415d-845c-7fab3f7e4763 · outbound

This paper cites The Curious Case of Neural Text Degeneration.

Statistical Multicriteria Evaluation of LLM-Generated Text The Curious Case of Neural Text Degeneration

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.755057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.755057Z digest=sha256:e384a06fc7933ff472e6a1f8452a822871bc14538f6be8ce2df29eb29574298c

Observation 3eb4cdf9-2109-46e9-8d61-42e0b965f71d · outbound

This paper cites Hothorn, F.

Statistical Multicriteria Evaluation of LLM-Generated Text Hothorn, F

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.250826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.760019Z digest=sha256:58ce133629a06287997156eab18a0edccf846fd6e6c7b558b76e9875cc21e3aa

Observation 482aa6f5-0bbe-4eb1-950c-6434c19dd408 · outbound

This paper cites Open graph benchmark: Datasets for machine learning on graphs.

Statistical Multicriteria Evaluation of LLM-Generated Text Open graph benchmark: Datasets for machine learning on graphs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.764669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.764669Z digest=sha256:40dd4f47c330818b0ccf1fc74b07cd519139da649b1f2f4d7932db39673ca473

Observation d4c2c78a-1f2f-4702-83eb-305930651da5 · outbound

This paper cites an unresolved cited work.

Statistical Multicriteria Evaluation of LLM-Generated Text Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:01:24.225154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.768921Z digest=sha256:0ef4131bdb57feba4775e5187e20c865537daff4cfb98fe8d6a803bf5d4fd35d

Observation d6245343-16f7-468b-be3d-58c25aa7a060 · outbound

This paper cites Jansen, G.

Statistical Multicriteria Evaluation of LLM-Generated Text Jansen, G

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.208171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.773521Z digest=sha256:514c4afe1f350bf999763c37fef858dfd5934d58697af240668735ccc66a8330

Observation 7f5a2312-0e66-4ddb-a436-6c2159c36e4e · outbound

This paper cites Jansen, H.

Statistical Multicriteria Evaluation of LLM-Generated Text Jansen, H

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.191725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.777980Z digest=sha256:f6737639ede5e878d53ab18cac9d470f2b0ea63beb9f2d9a9ac0f78dfeb92fde

Observation abde3ebb-d0fc-4b16-872b-7c4b2e5a2439 · outbound

This paper cites Contributions to the Decision Theoretic Foundations of Machine Learning and Robust Statistics under Weakly Structured Information.

Statistical Multicriteria Evaluation of LLM-Generated Text Contributions to the Decision Theoretic Foundations of Machine Learning and Robust Statistics under Weakly Structured Information

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.782509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.782509Z digest=sha256:e979ce71847239dc8fed1d073f7ed920da964e683cd4ac04b9752274d63ce77d

Observation 0d9394e2-1fc2-49d7-a269-ea2bb814f8ad · outbound

This paper cites Statistical comparisons of classifiers by generalized stochastic dominance.

Statistical Multicriteria Evaluation of LLM-Generated Text Statistical comparisons of classifiers by generalized stochastic dominance

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.175823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.787435Z digest=sha256:f626a94c48bfcaf55cc5ac4e5d9a6f2277212c691cef181529b98e30ad6921ac

Observation ee6a52ff-a88f-4ff5-9a52-0ef5ac3231dc · outbound

This paper cites Multi-target decision making under conditions of severe uncertainty.

Statistical Multicriteria Evaluation of LLM-Generated Text Multi-target decision making under conditions of severe uncertainty

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.158971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.792086Z digest=sha256:dd9deb18a674d911816e75ebec2005703ff3134eca6a15af47811ea3172cc5d6

Observation 9478c78a-6f67-461b-8abc-620cc3c3064c · outbound

This paper cites Robust statistical comparison of random variables with locally varying scale of measurement.

Statistical Multicriteria Evaluation of LLM-Generated Text Robust statistical comparison of random variables with locally varying scale of measurement

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.143662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.796599Z digest=sha256:d7be1f4ecaad484d5d3dff4e4903b6a1d6bdf7a113551829233dc713283bbc37

Observation 2e7b3895-f8a9-44d8-9276-bcecca37765b · outbound

This paper cites Statistical multicriteria benchmarking via the GSD -front.

Statistical Multicriteria Evaluation of LLM-Generated Text Statistical multicriteria benchmarking via the GSD -front

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.126866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.801334Z digest=sha256:fa1406c5682ed0c5a329480011906ab9e78487c8bd03fa01409fa4ff05373231

Observation 47d7e42a-b273-4d21-9b91-4c51e05154ad · outbound

This paper cites Jelinek, R.

Statistical Multicriteria Evaluation of LLM-Generated Text Jelinek, R

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.805751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.805751Z digest=sha256:d90fe1825247d44ddd4270df6e8efb36fdae4a54591f416f3577dcb19e852848

Observation 405d4636-93bb-47fd-989a-03089d3a33a9 · outbound

This paper cites The prism alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models.

Statistical Multicriteria Evaluation of LLM-Generated Text The prism alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.106718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.810810Z digest=sha256:fe669070b1241cc310dbefc6d42ab7ab71fa0f40a631e235e05cd761fbecda45

Observation 552b9996-ea13-499c-9877-1c36bc3abfd1 · outbound

This paper cites Efficient multi-criteria optimization on noisy machine learning problems.

Statistical Multicriteria Evaluation of LLM-Generated Text Efficient multi-criteria optimization on noisy machine learning problems

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.084317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.815635Z digest=sha256:0aeac947fa4304d54169910878c5b0e51211f3f42cc58c45d40268616d331f69

Observation ad91afc2-cd04-4249-8be3-23b5cd9e65f1 · outbound

This paper cites Towards quantifying the effect of datasets for benchmarking: A look at tabular machine learning.

Statistical Multicriteria Evaluation of LLM-Generated Text Towards quantifying the effect of datasets for benchmarking: A look at tabular machine learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.068837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.820250Z digest=sha256:96ca689e6f2aeabf9dfd0c4f0914c69834dbd9807366461bf13e90baa0fcdfc6

Observation 7e220059-a974-498a-bdc9-bd69afd6e9f9 · outbound

This paper cites Accelerated experimental design using a human-ai teaming framework.

Statistical Multicriteria Evaluation of LLM-Generated Text Accelerated experimental design using a human-ai teaming framework

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.051460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.825117Z digest=sha256:90b3236e69028aec164e30a414286312b900b2a5f53dc41f9b22f50b76cb43de

Observation 8363a457-cb5f-41f4-8e89-a1f6dc33daab · outbound

This paper cites Factuality enhanced language models for open-ended text generation.

Statistical Multicriteria Evaluation of LLM-Generated Text Factuality enhanced language models for open-ended text generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.030793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.829565Z digest=sha256:9a926383f39a0821903d530c91060808a35773ffea7c2412215215b0f408033d

Observation 94544a59-e947-480a-8e71-1e8d656d5c9e · outbound

This paper cites A diversity-promoting objective function for neural conversation models.

Statistical Multicriteria Evaluation of LLM-Generated Text A diversity-promoting objective function for neural conversation models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.010099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.834159Z digest=sha256:34c89486e00ab73934eabf33fe20aef29fe8be8aaa1167411c3e95c47e0fd9be

Observation 57e6a772-cc2b-407d-b354-7fa7a9464712 · outbound

This paper cites Contrastive decoding: Open-ended text generation as optimization, 2023.

Statistical Multicriteria Evaluation of LLM-Generated Text Contrastive decoding: Open-ended text generation as optimization, 2023

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.839032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.839032Z digest=sha256:0fbe61ce52e38f47c6a5a57e748844d3918f6d88a5254622e26ef6e561adc772

Observation 875b34cc-8ac6-4733-9fe6-ab580adcc110 · outbound

This paper cites Quantifying Variance in Evaluation Benchmarks.

Statistical Multicriteria Evaluation of LLM-Generated Text Quantifying Variance in Evaluation Benchmarks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.843911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.843911Z digest=sha256:268de2b3e4bdd3e8fce4057656618cc8369e5a960fe4067e2a18bd7e1b667b81

Observation 8993ba60-1552-4443-8894-7bb642e477e4 · outbound

This paper cites Pointer sentinel mixture models, 2016.

Statistical Multicriteria Evaluation of LLM-Generated Text Pointer sentinel mixture models, 2016

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.849573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.849573Z digest=sha256:9b9ac3291d81dffc4934f7b7dd0e87ee31e75683e10b8de777c3b245ee9ece98

Observation cecee8ff-15b0-4016-a5c2-1b9d7fb8f91e · outbound

This paper cites Mersmann, M.

Statistical Multicriteria Evaluation of LLM-Generated Text Mersmann, M

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.967599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.854610Z digest=sha256:f106dfaf834a7d7c6e15c104ca32dd4b2a2a7e4f72eb8cf1b28e481e319ed58e

Observation f3b8f4af-bcea-4658-8d44-d31cd6c37a8b · outbound

This paper cites Meyer, F.

Statistical Multicriteria Evaluation of LLM-Generated Text Meyer, F

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.953009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.859368Z digest=sha256:8d2b2c4c5d08b3759ab3b91fa8c61a33b18fcca8d51a9f8bc24fee5c7684404d

Observation 006789af-20db-4f85-9519-c68d78d8394c · outbound

This paper cites Learning de-biased regression trees and forests from complex samples.

Statistical Multicriteria Evaluation of LLM-Generated Text Learning de-biased regression trees and forests from complex samples

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.937613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.863706Z digest=sha256:d1b067339f03adcf1c696c647c4015ac4620379f79d4a54f4ff302952ca14eac

Observation d32f776c-4f0f-494a-b456-b7912e58b482 · outbound

This paper cites an unresolved cited work.

Statistical Multicriteria Evaluation of LLM-Generated Text Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:01:23.921306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.868285Z digest=sha256:d302e15f6de947a0705b3a1d20b8ea638ff91acd45b37df1a98e71503bc10d2e

Observation 5fd61fd8-57aa-4927-b217-4cdf03e943ec · outbound

This paper cites Mauve: Measuring the gap between neural text and human text using divergence frontiers.

Statistical Multicriteria Evaluation of LLM-Generated Text Mauve: Measuring the gap between neural text and human text using divergence frontiers

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.905053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.872613Z digest=sha256:2d26e5cf2c1083d8f83f151740afb4ccf0a1b819ec7b0b9a2e3501c8fe145d83

Observation 30056779-8cec-45f2-b62c-1190e981c105 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Statistical Multicriteria Evaluation of LLM-Generated Text Direct preference optimization: Your language model is secretly a reward model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.876746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.876746Z digest=sha256:d3e4193f070716604421657a21fd5d8513efce3bc05cda9046d5acc8b6c5b749

Observation dd83a02d-229d-417e-a2c9-46d02366bcf8 · outbound

This paper cites Partial rankings of optimizers.

Statistical Multicriteria Evaluation of LLM-Generated Text Partial rankings of optimizers

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.879394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.881040Z digest=sha256:4d2bd4be49533c3b8b723f22506fadb78f136f05589039b91e8ac0419731b642

Observation c5a11e91-667c-4885-b3d2-382221ab37b1 · outbound

This paper cites Levelwise data disambiguation by cautious superset classification.

Statistical Multicriteria Evaluation of LLM-Generated Text Levelwise data disambiguation by cautious superset classification

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.862698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.885608Z digest=sha256:3860b71e905e0875f7e9bd9b920ca2be116b832dc239ee2d2730ed92a1a72dcf

Observation e07be0ab-dc50-4d13-9a55-4209ff7d99e0 · outbound

This paper cites Approximately bayes-optimal pseudo-label selection.

Statistical Multicriteria Evaluation of LLM-Generated Text Approximately bayes-optimal pseudo-label selection

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.846974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.890326Z digest=sha256:daaa207dde4f7e472c834267da182d06636d01b591fbc55d8d599f4c4b0f1efe

Observation 75b9fd2e-c80a-47e0-abcb-a0e83b233fd3 · outbound

This paper cites In all likelihoods: Robust selection of pseudo-labeled data.

Statistical Multicriteria Evaluation of LLM-Generated Text In all likelihoods: Robust selection of pseudo-labeled data

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.830638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.894647Z digest=sha256:dc157ce4ebb5a8b176c30e7262cf7392a85a20e709d5181011d0e07fd26ee0aa

Observation 5b254428-beea-4c10-9932-a0170c11af7b · outbound

This paper cites Explaining Bayesian Optimization by Shapley Values Facilitates Human-AI Collaboration.

Statistical Multicriteria Evaluation of LLM-Generated Text Explaining Bayesian Optimization by Shapley Values Facilitates Human-AI Collaboration

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:01:23.280879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.899482Z digest=sha256:b22cff0b5dd965d94cea901ed96cdaccecd171cc33295151aac437e862054c8b

Observation 7053843b-859e-4f0e-bdc4-214a69dae1a9 · outbound

This paper cites A Statistical Case Against Empirical Human-AI Alignment.

Statistical Multicriteria Evaluation of LLM-Generated Text A Statistical Case Against Empirical Human-AI Alignment

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:01:23.257694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.904326Z digest=sha256:819b7bba69929f170990ada9539070b7fa0ceb12f5a548743c8f8976d8650efa

Observation e1478825-dbcd-4ec9-86c8-772f9d969d6d · outbound

This paper cites A meta-analysis of overfitting in machine learning.

Statistical Multicriteria Evaluation of LLM-Generated Text A meta-analysis of overfitting in machine learning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.815993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.909152Z digest=sha256:04dab00b6148a6b3ba5261b6f53d6e6ecf25f28acc9f74906821e5606d7ea13f

Observation 8c2823b6-54a7-4337-b4fc-06440b25f108 · outbound

This paper cites Schneider, L.

Statistical Multicriteria Evaluation of LLM-Generated Text Schneider, L

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.800614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.913529Z digest=sha256:dec32e1ffaa940abeb317b8ba482ccd91c62449ee5d1dc56d7906dd4b512dfbf

Observation 96a68390-0e70-4846-8512-f23487770614 · outbound

This paper cites A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes.

Statistical Multicriteria Evaluation of LLM-Generated Text A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.781777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.918228Z digest=sha256:dd0dc68843934617410747f0fe2d6166bde9e6a381342d16612355ecabe1ee64

Observation c232a105-626a-44e4-a2d6-a8d84a0dc1c4 · outbound

This paper cites Shirali, R.

Statistical Multicriteria Evaluation of LLM-Generated Text Shirali, R

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.764739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.922919Z digest=sha256:7df0b003b33743869ba137d9491fb23804af49686b7cec23be541cbad378886b

Observation e0938fbb-fa31-4e29-95c0-98b183dd0140 · outbound

This paper cites An empirical study on contrastive search and contrastive decoding for open-ended text generation, 2022.

Statistical Multicriteria Evaluation of LLM-Generated Text An empirical study on contrastive search and contrastive decoding for open-ended text generation, 2022

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.746915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.927214Z digest=sha256:af39f1e849d6601b9c83b6ed4e30372d6b6a0d8c0b6784dcdcd947e919e25eb0

Observation 8aa759bd-0119-472c-9ecb-b54d50bca22e · outbound

This paper cites A contrastive framework for neural text generation, 2022.

Statistical Multicriteria Evaluation of LLM-Generated Text A contrastive framework for neural text generation, 2022

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.730152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.931636Z digest=sha256:8d9dc06e32a5473bffc1b93f76302bd5f229f6ebd3373855ff85b73d6c5f6ad3

Observation 8d2b3088-51e4-4876-91a8-b01909cd9031 · outbound

This paper cites Evaluating the evaluation of diversity in natural language generation.

Statistical Multicriteria Evaluation of LLM-Generated Text Evaluating the evaluation of diversity in natural language generation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.713796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.935910Z digest=sha256:4f07c32e75e618fa790befd2498b112f12056efdbdeaf9e25442b43e7e8a51ba

Observation 3f3b1e7d-5e75-4367-bc58-f7eac5178bc6 · outbound

This paper cites Scientific machine learning benchmarks.

Statistical Multicriteria Evaluation of LLM-Generated Text Scientific machine learning benchmarks

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.696300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.940493Z digest=sha256:1d25930e1f5194c25113b2f90733b0b57765477f2532c9e0a2f45b878e790ee1

Observation 25df4b8e-684c-49ce-ae01-76e76555ceab · outbound

This paper cites Openml: networked science in machine learning.

Statistical Multicriteria Evaluation of LLM-Generated Text Openml: networked science in machine learning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.944941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.944941Z digest=sha256:65912b83efbd4879a4c1a24b033758e9c8f4729e4dafce3efbbeca014c7cc64b

Observation 688c6c05-f339-4aca-aea3-74751283f638 · outbound

This paper cites an unresolved cited work.

Statistical Multicriteria Evaluation of LLM-Generated Text Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.949512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.949512Z digest=sha256:8f8f4e59ad0a87bdf404e714f252d98434612a03a5bf4c748bde04121dffc556

Observation 69419229-b8ae-4126-ae63-cd1742e20623 · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

Statistical Multicriteria Evaluation of LLM-Generated Text LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.954002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.954002Z digest=sha256:75e9289263fab1728a8d7eb964b1630d1fadf651343ddcc2e5def07f9f480c8b

Observation 79b025d7-184c-4c3c-8d74-3b23bac0e1bd · outbound

This paper cites Principled Bayesian Optimisation in Collaboration with Human Experts.

Statistical Multicriteria Evaluation of LLM-Generated Text Principled Bayesian Optimisation in Collaboration with Human Experts

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.958957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.958957Z digest=sha256:57d0bf5f63d72d6c2c4cd1de362a29a7324b3f8cc414b2bc648b4d3fda5b4099

Observation 02fad94f-8661-4f1e-a412-cf73f8c6bb4f · outbound

This paper cites Qwen2 Technical Report.

Statistical Multicriteria Evaluation of LLM-Generated Text Qwen2 Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.964167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.964167Z digest=sha256:8e32e334d62e9c17ed784977f40e4050622df2ff0d539d50a3958d428550f896

Observation 9dd85e22-0986-42fd-9dd0-adf64cea4b89 · outbound

This paper cites Benchmarking llms via uncertainty quantification.

Statistical Multicriteria Evaluation of LLM-Generated Text Benchmarking llms via uncertainty quantification

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.657388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.969120Z digest=sha256:f17381f2944805607b67406d2daafc1cdd201d84d57e8f0bac489c3ca70a00c7

Observation 80df0dbc-876b-48e6-b825-c1c79dbccd9c · outbound

This paper cites Zhang and M.

Statistical Multicriteria Evaluation of LLM-Generated Text Zhang and M

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.641910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.973816Z digest=sha256:d977e1a0c12a80eaa2c3e04f3bfa519b864e4198021d0cf63eb1901d1263d90f

Observation dee8534b-81d3-4281-b9c6-c88936471016 · outbound

This paper cites Inherent trade-offs between diversity and stability in multi-task benchmarks.

Statistical Multicriteria Evaluation of LLM-Generated Text Inherent trade-offs between diversity and stability in multi-task benchmarks

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.626263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.978485Z digest=sha256:83ab48fe00bdadd8ec9895aace8c17ef4269ff5c8edc3611841d4101da91df9e

Observation dcc7d216-632a-4d04-98b1-03c0e069dab1 · outbound

This paper cites Zhang, M.

Statistical Multicriteria Evaluation of LLM-Generated Text Zhang, M

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.607704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.982795Z digest=sha256:daec399b686c1a0aa97ff75285819d455262a7d953faaaecc749cf646268ba42

Observation bbbe07c5-3379-493f-b2b2-4478d742f342 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Statistical Multicriteria Evaluation of LLM-Generated Text OPT: Open Pre-trained Transformer Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.987507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.987507Z digest=sha256:1507dd58a2b3eaa0998c983d75a8b6557ec7e6ff29d8124b9626158c23194cf3

Observation ce165ab0-9b6b-4fd7-b5c3-b509f1e98262 · outbound

This paper cites Don't Make Your LLM an Evaluation Benchmark Cheater.

Statistical Multicriteria Evaluation of LLM-Generated Text Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.992886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.992886Z digest=sha256:dfcad7d1b53b682abad591c84742ad2099d3e88f8a9fa1ab280350b59a034a95

Observation 84ec1c5f-a10f-4a70-aa16-4e7972dd720e · outbound

This paper cites Time-Varying Gaussian Process Bandits with Unknown Prior.

Statistical Multicriteria Evaluation of LLM-Generated Text Time-Varying Gaussian Process Bandits with Unknown Prior

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:01:23.127948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.998385Z digest=sha256:73a34252f7f6f9880e297c1bfb7eef487bba8459cf7620579b8da72ef23afd4f

Observation f055e5b2-a437-4db0-beb4-bd59ff17b15d · outbound

This paper cites write newline.

Statistical Multicriteria Evaluation of LLM-Generated Text write newline

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:23.004618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:23.004618Z digest=sha256:afb3d7d9b6caa440b7d4ea3a0627aca51c875c5bfeb0a51b405c769589cced52

Observation 937452c5-e766-4b61-a551-55a656f248d4 · outbound

This paper cites @esa (Ref.

Statistical Multicriteria Evaluation of LLM-Generated Text @esa (Ref

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:23.010839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:23.010839Z digest=sha256:02fe8de0ce6712891518ee48311d888fe41ac0fe641994d09cb9793e8238a9b6

Observation 85b98448-6d2f-48b0-85f8-7259e64911d8 · outbound

This paper cites an unresolved cited work.

Statistical Multicriteria Evaluation of LLM-Generated Text Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:23.017499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:23.017499Z digest=sha256:92e8ee8d76ed20dfd2bab16cc9122b4929aaa61178a7da04aa15088423681e6a

Observation 07c14cbb-1a00-47ed-a09a-6fc1bb80fc65 · outbound

This paper cites an unresolved cited work.

Statistical Multicriteria Evaluation of LLM-Generated Text Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:23.023749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:23.023749Z digest=sha256:8f2e2723aaee562902e55e1b78092d7614ab8d88c721db8f4fc97d803c4b73ce

Pith citing papers

Observation 9b8b12fd-8aec-424e-83be-138f89dee1bc · inbound

Self-Supervised Laplace Approximation for Bayesian Uncertainty Quantification cites this paper.

Self-Supervised Laplace Approximation for Bayesian Uncertainty Quantification Statistical Multicriteria Evaluation of LLM-Generated Text

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:52:17.137374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-13T04:49:15.239636Z digest=sha256:7e373c0875e7027b68b19be3af5371c4a62312b0a61f9d508dd952c9a7a1086e