Pith. sign in

Paper Citation Record · LEDGER

A Framework for Evaluating LLMs Under Task Indeterminacy

As of 12 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2411.13760.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13760 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:58:47.648449Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact2
  • verified fuzzy16
  • unresolved24
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c077c67-c093-4ebf-88d3-ecb29673ac03 · outbound

This paper cites DICES Dataset: Diversity in Conversational AI Evaluation for Safety.

A Framework for Evaluating LLMs Under Task Indeterminacy DICES Dataset: Diversity in Conversational AI Evaluation for Safety

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.839258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.410920Z digest=sha256:8d67c756dbfeb130689afa90084c3d9278792bd406995c0a2e3dad2997100a25

Observation a57fd76d-a59b-4874-9ec1-70b47786c546 · outbound

This paper cites Stop Measuring Calibration When Humans Disagree.

A Framework for Evaluating LLMs Under Task Indeterminacy Stop Measuring Calibration When Humans Disagree

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.416548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.416548Z digest=sha256:8d30c8256a0e53b5dea633ec34cdbf8f8d63109b69f74b702c3eced10f760683

Observation 2b3da26a-e236-43ed-9e86-cac277deb601 · outbound

This paper cites It’s the End of the Gold Standard as we Know it.

A Framework for Evaluating LLMs Under Task Indeterminacy It’s the End of the Gold Standard as we Know it

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.821568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.422176Z digest=sha256:a0b09546be6f764e80bcdb6900dc5b20a68c19aae1af85fa491539889a8e1ec3

Observation 8d37a0ce-5096-4610-a17a-d9fba6cdd5c2 · outbound

This paper cites Like trainer, like bot? Inheritance of bias in algorithmic content moderation.

A Framework for Evaluating LLMs Under Task Indeterminacy Like trainer, like bot? Inheritance of bias in algorithmic content moderation

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T15:58:48.520699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.427557Z digest=sha256:19c65030c8ba7886632d0536dfde786fb85ad3d8cdfb42ca283a4ad526b79ddf

Observation 28c48695-e4b3-47dc-8a24-ac5d9e0ce8b6 · outbound

This paper cites Holistic evaluation of language models.

A Framework for Evaluating LLMs Under Task Indeterminacy Holistic evaluation of language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.802595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.433122Z digest=sha256:6c10fdca16b8cb59669dc3377e5b5805b8ab75b7b54b5ad88c8409924e9051c3

Observation 28671047-f62c-4e76-bf92-a379645b173b · outbound

This paper cites an unresolved cited work.

A Framework for Evaluating LLMs Under Task Indeterminacy Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.438380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.438380Z digest=sha256:7d96710c472705b429f7adedcf3e5f1cbeade4db08134d0d9f6f71e88df1b46c

Observation 00d71071-c431-4e74-8027-1c47e9925021 · outbound

This paper cites Revolt: Collaborative Crowdsourcing for Labeling Machine Learning Datasets.

A Framework for Evaluating LLMs Under Task Indeterminacy Revolt: Collaborative Crowdsourcing for Labeling Machine Learning Datasets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.447529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.447529Z digest=sha256:ca87952d8fb033728c5de64fb19c0d95aea1a1bd913ff8b6aa4aab5d3750344a

Observation c6af6ccc-9894-468b-b70a-b43f38dc4034 · outbound

This paper cites Judgment sieve: Reducing uncertainty in group judgments through interventions targeting ambiguity versus disagreement.

A Framework for Evaluating LLMs Under Task Indeterminacy Judgment sieve: Reducing uncertainty in group judgments through interventions targeting ambiguity versus disagreement

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.786308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.453452Z digest=sha256:94d51b4cb627f90a0e06f2ac9c7cd5cc9910643dfd7b5613e709309416176b70

Observation 9a458d73-d10a-4790-80a3-159d996761a7 · outbound

This paper cites Judgment Sieve: Reducing Uncertainty in Group Judgments through Interventions Targeting Ambiguity versus Disagreement.

A Framework for Evaluating LLMs Under Task Indeterminacy Judgment Sieve: Reducing Uncertainty in Group Judgments through Interventions Targeting Ambiguity versus Disagreement

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:58:48.356776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.459514Z digest=sha256:663b07087cab277d99e35559a79c388039cb0295b7de2e377679899ee2bfe938

Observation 9b8ed86a-925b-4376-959a-e23aace9ef7d · outbound

This paper cites Dealing with Disagree- ments: Looking Beyond the Majority V ote in Subjective Annotations.

A Framework for Evaluating LLMs Under Task Indeterminacy Dealing with Disagree- ments: Looking Beyond the Majority V ote in Subjective Annotations

Reference 10

Resolution
malformed identifier
doi_truncated, observed 2026-08-12T15:58:47.793619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.465135Z digest=sha256:9c851f70da182c5a52491c78f3292b6a7aa1a0bdde76a06c6ddc1e4a2a87abcf

Observation ff854f86-582a-476c-b805-f332c7290525 · outbound

This paper cites D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and Evaluation.

A Framework for Evaluating LLMs Under Task Indeterminacy D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and Evaluation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.470355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.470355Z digest=sha256:497df2fff16ede0e32947d5a811b30c7572a47637695f301f7aeaeac6d1d1639

Observation 3ea58abb-d693-4056-aac1-58b585a5ad0a · outbound

This paper cites Red-Teaming for Generative AI: Silver Bullet or Security Theater?.

A Framework for Evaluating LLMs Under Task Indeterminacy Red-Teaming for Generative AI: Silver Bullet or Security Theater?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.475898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.475898Z digest=sha256:7b92f3a8d24ad03841c868c1b88a37643101dac0ef9dccf9a8f35fab40c3c274

Observation 8aed3702-9d7b-445e-aefd-781624d01b0f · outbound

This paper cites Efficient Conformal Prediction via Cascaded Inference with Expanded Admission.

A Framework for Evaluating LLMs Under Task Indeterminacy Efficient Conformal Prediction via Cascaded Inference with Expanded Admission

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.481439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.481439Z digest=sha256:16464f3f7d8ba86f82ff99c0d681d71b54d866e3d7a178771743357f52930e08

Observation 0ea16206-678b-402d-8ee8-379a7cd8505f · outbound

This paper cites The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels.

A Framework for Evaluating LLMs Under Task Indeterminacy The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.486664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.486664Z digest=sha256:d8feba310852da0fb4030370208fdd82de7e9fb570301860b0eb5314f963b1ce

Observation cb126527-a912-4f1a-a3e6-558a7fc321a7 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

A Framework for Evaluating LLMs Under Task Indeterminacy Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.491762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.491762Z digest=sha256:c593d14c3f67193121fb6f05988fe4f9e72016d1254ea1c3003a4cbc5e5f5f12

Observation 68160aa2-5850-4122-90bc-0f1ae0396ae7 · outbound

This paper cites Deep Label Distribution Learning With Label Ambiguity.

A Framework for Evaluating LLMs Under Task Indeterminacy Deep Label Distribution Learning With Label Ambiguity

Reference 16

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T15:58:48.243981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.496224Z digest=sha256:a8445c43a36627d6fada49eb294010c3cbc25a3b5809215740fe762f0424db68

Observation 4809e15c-34b0-4d78-b633-a8e1ed9dc364 · outbound

This paper cites Label Distribution Learning.

A Framework for Evaluating LLMs Under Task Indeterminacy Label Distribution Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.500370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.500370Z digest=sha256:7abdd97c7be9dacdb6ced5422846f15a392b79f895eeb7a4c0008025c0796dcd

Observation 210bc53a-7606-4252-9fbc-38252d34451d · outbound

This paper cites The disagreement deconvolution: Bringing machine learning performance metrics in line with reality.

A Framework for Evaluating LLMs Under Task Indeterminacy The disagreement deconvolution: Bringing machine learning performance metrics in line with reality

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.769904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.504721Z digest=sha256:7c081d4a452488cbfd56af245dc992befefd153d9ea2839513ead5a382565880

Observation a7959e7b-177f-4f5d-a252-6e4a117e8c02 · outbound

This paper cites Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S.

A Framework for Evaluating LLMs Under Task Indeterminacy Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.508867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.508867Z digest=sha256:9486cf2e5b8043f57803617ae5620d69bbf2d716f1863c637fc1656896d86874

Observation a7d6d6bd-d8b1-45c2-9a75-a31e0516dfad · outbound

This paper cites Jury learning: Integrating dissenting voices into machine learning models.

A Framework for Evaluating LLMs Under Task Indeterminacy Jury learning: Integrating dissenting voices into machine learning models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.752362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.513299Z digest=sha256:e247660ff881d665f3805bdc0328de2bd3cb343b0960255decf96db6ecb861ca

Observation 9cbcbbe0-069c-473b-9568-f8cd5de417a1 · outbound

This paper cites Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation.

A Framework for Evaluating LLMs Under Task Indeterminacy Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.734640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.517490Z digest=sha256:3e9ba7e3d5023f8acf57232f9314e7c708d4cf99c3db37a11790e33eb7efaa42

Observation 50d372cd-4d44-4a5d-aa16-ca30d7b7bb7b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

A Framework for Evaluating LLMs Under Task Indeterminacy Measuring Massive Multitask Language Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.522165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.522165Z digest=sha256:9448e1a26a4161ae9f53052b069ff9699fa850cd3bb1f326aa5ce6b34ae75d66

Observation 28b35d4a-d867-48bc-9e37-4ffe9a43f1d8 · outbound

This paper cites Intersectionality in ai safety: Using multilevel models to understand diverse perceptions of safety in conversational ai.

A Framework for Evaluating LLMs Under Task Indeterminacy Intersectionality in ai safety: Using multilevel models to understand diverse perceptions of safety in conversational ai

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.717943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.527335Z digest=sha256:6802010f4dd3428413bc476b7b00c4ec7810e9cb8a2d68492ecf5ede189c1133

Observation 1bf6244c-5633-479d-990b-713b16189533 · outbound

This paper cites Culturally Aware Natural Language Inference.

A Framework for Evaluating LLMs Under Task Indeterminacy Culturally Aware Natural Language Inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.531722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.531722Z digest=sha256:d4db82c73c1b445bf59d8950c82901b96a4145263fa6acda4be48354a7462a7a

Observation 77dd0d86-8374-4649-9493-e5226e7d39dd · outbound

This paper cites Chaithanya Manam, Dwarakanath Jampani, Mariam Zaim, Meng-Han Wu, and Alexander J.

A Framework for Evaluating LLMs Under Task Indeterminacy Chaithanya Manam, Dwarakanath Jampani, Mariam Zaim, Meng-Han Wu, and Alexander J

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.699453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.536697Z digest=sha256:35c115edba22e707a03be5f3fbe846af5cc1c38bfdbaa097e74dccbb666f5041

Observation a874b2e7-e4c3-487c-88ed-ec1d04b97ef8 · outbound

This paper cites Annotation Error Detection: Analyzing the Past and Present for a More Coherent Future.

A Framework for Evaluating LLMs Under Task Indeterminacy Annotation Error Detection: Analyzing the Past and Present for a More Coherent Future

Reference 26

Resolution
malformed identifier
no resolver link, observed 2026-08-12T15:58:47.541649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.541649Z digest=sha256:e4fa32845081b1fb1d433070684762030e99e5a4793822010d0fd9f7282e9022

Observation d383fbc1-768b-4684-90d8-75a83dca527e · outbound

This paper cites A Bayesian Framework for Modeling Human Evaluations.

A Framework for Evaluating LLMs Under Task Indeterminacy A Bayesian Framework for Modeling Human Evaluations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.546651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.546651Z digest=sha256:c9a0749ee8bfe0ae4e05e1f28802ba688eab3d032953fd5f3815d6ef5bd278a2

Observation d3ef8992-3d96-46b0-ad4e-d464487125d0 · outbound

This paper cites Reconsidering Annotator Disagreement about Racist Language: Noise or Signal?.

A Framework for Evaluating LLMs Under Task Indeterminacy Reconsidering Annotator Disagreement about Racist Language: Noise or Signal?

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.679755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.552442Z digest=sha256:bbf80840df22df60175894a2e897b0f49e6fb2089334142ec55ae2bb493690f3

Observation c16a2fd9-fa2f-488f-aada-166b24f23f4d · outbound

This paper cites Learning to predict population-level label distributions.

A Framework for Evaluating LLMs Under Task Indeterminacy Learning to predict population-level label distributions

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.661299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.557193Z digest=sha256:c9d0a0227ab927e07f1f9e130c5778dcccbc8505e684f9d3c9a0394703bb9422

Observation f5c68c08-3d94-4613-8b4f-72dae82cb0c8 · outbound

This paper cites A Safe Harbor for AI Evaluation and Red Teaming.

A Framework for Evaluating LLMs Under Task Indeterminacy A Safe Harbor for AI Evaluation and Red Teaming

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.561742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.561742Z digest=sha256:3ce2147fb26afd1ee7a185287a39bbfe6a5646c568cfc812488655bf5435df19

Observation b032cac8-7e8e-4f0f-ac4b-c349810d2c66 · outbound

This paper cites A Framework for Automated Measurement of Responsible AI Harms in Generative AI Applications.

A Framework for Evaluating LLMs Under Task Indeterminacy A Framework for Automated Measurement of Responsible AI Harms in Generative AI Applications

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.566997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.566997Z digest=sha256:3aeef417ccdff7e238ff05d0f2342ca9a8e9174b3e424735d9622f7091201188

Observation e9a420e2-695e-46aa-97db-9bd5713d93c7 · outbound

This paper cites an unresolved cited work.

A Framework for Evaluating LLMs Under Task Indeterminacy Unresolved cited work

Reference 32

Resolution
verified exact
doi, observed 2026-08-12T15:58:47.716100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.572311Z digest=sha256:b87a5117b9cd65801485d2eee0128d31febb8dfcfb90de5efd75939cd6a93506

Observation 630a4f9b-ddfe-4f13-bd99-180ec7bc36f4 · outbound

This paper cites StereoSet: Measuring stereotypical bias in pretrained language models.

A Framework for Evaluating LLMs Under Task Indeterminacy StereoSet: Measuring stereotypical bias in pretrained language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.577191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.577191Z digest=sha256:cfe1cf243f8e8095c3874b50ad886c083c2a722a2eb5f1ba2c8460827d0352b3

Observation 6657a046-e9a1-4318-a733-e6ca4ee0f37e · outbound

This paper cites Diversity-aware annotation for conversational ai safety.

A Framework for Evaluating LLMs Under Task Indeterminacy Diversity-aware annotation for conversational ai safety

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.645028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.581750Z digest=sha256:3fc90fa98d49938629eaccdd6cf1ddfe6f809d0fd458ea467fa37479069b5962

Observation a6e9d212-c5aa-486e-ac2b-6bb52305a3f3 · outbound

This paper cites Inherent disagreements in human textual inferences.

A Framework for Evaluating LLMs Under Task Indeterminacy Inherent disagreements in human textual inferences

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.627575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.586220Z digest=sha256:73854e0eb65c3f223d2314fdc32ee9a0f546266284750dca4910517182988c14

Observation 4540f88d-e957-4faf-8bf6-61f13e4b7019 · outbound

This paper cites Inherent Disagreements in Human Textual Inferences.

A Framework for Evaluating LLMs Under Task Indeterminacy Inherent Disagreements in Human Textual Inferences

Reference 36

Resolution
malformed identifier
no resolver link, observed 2026-08-12T15:58:47.591125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.591125Z digest=sha256:31d38a26b3de58c723f73e9e091a1bd1a9e5a638b11c6dc3054feec6a2661f54

Observation ea6b370c-3186-416d-83f1-dc46594621fe · outbound

This paper cites Human uncertainty makes classification more robust.

A Framework for Evaluating LLMs Under Task Indeterminacy Human uncertainty makes classification more robust

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.610984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.595776Z digest=sha256:d23a5691d0015b909cb416eb2ecad753572d1db809569e88e86cdc8bf5924fbc

Observation 8c67e5e5-81c0-44dc-87a4-d46a433706d4 · outbound

This paper cites The 'Problem' of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation.

A Framework for Evaluating LLMs Under Task Indeterminacy The 'Problem' of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.600485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.600485Z digest=sha256:56fbae162546a3a758695eda8af271f24d5f9821bc79e206d30b9399b73d064f

Observation 272987e6-d92a-44e6-a2a4-c316535e7c9a · outbound

This paper cites On Releasing Annotator-Level Labels and Information in Datasets.

A Framework for Evaluating LLMs Under Task Indeterminacy On Releasing Annotator-Level Labels and Information in Datasets

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.605264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.605264Z digest=sha256:b770e9cd1bf258d9daa57188e1423bf76b3b2460c23b36626de3187746b74348

Observation f1e7394c-d3f0-43e8-b506-d192bc39d250 · outbound

This paper cites Principled Evaluation with Human Labels: One Rater at a Time and Rater Equivalence.

A Framework for Evaluating LLMs Under Task Indeterminacy Principled Evaluation with Human Labels: One Rater at a Time and Rater Equivalence

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T15:58:47.871195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.610350Z digest=sha256:cff163f88877215f1f6fdfcb1051c486bd35849ab9fae3f82f51066c1c23ca4c

Observation 602c5eaf-4359-4f96-8eaa-e3f5f29ae945 · outbound

This paper cites Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection.

A Framework for Evaluating LLMs Under Task Indeterminacy Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.615404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.615404Z digest=sha256:b5d79075e836b76af42c56d9635a8d5aabe4f67f5763d11e89e94f0d92e0548d

Observation 09088a7d-bbd7-4f5d-ac12-f8a906a710b3 · outbound

This paper cites Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences.

A Framework for Evaluating LLMs Under Task Indeterminacy Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.620627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.620627Z digest=sha256:cdd853af7353a2468d2879dca936242bb0547fc70a1faf4d4ac208cdf267ebc8

Observation 83efe431-4f75-4d3f-a57a-9139d956405b · outbound

This paper cites A case for soft loss functions.

A Framework for Evaluating LLMs Under Task Indeterminacy A case for soft loss functions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.625408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.625408Z digest=sha256:580b6d6d815a6bf2a6cd1d17edd23c83539fd02cef8695db1ec5aa0dc0129684

Observation e5c52f95-88a5-46ce-9d23-698238ed1636 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

A Framework for Evaluating LLMs Under Task Indeterminacy GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.630144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.630144Z digest=sha256:1e7238a875e62de4c89446cc00026b3e76b14cb844e3d1561ba28b7028c4fb34

Observation affa9a7a-e8f2-410a-a2c9-bcd63e415c00 · outbound

This paper cites gold data.

A Framework for Evaluating LLMs Under Task Indeterminacy gold data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.582824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.634597Z digest=sha256:4b123d93af1429e2b0449dfb351f025d11fc16e8ef08d6a4fd24ae5bbcb2d5e5

Observation 993a6c94-0d98-48a6-ac30-d1adeb6b00b7 · outbound

This paper cites Super- naturalinstructions:generalization via declarative instructions on 1600+ tasks.

A Framework for Evaluating LLMs Under Task Indeterminacy Super- naturalinstructions:generalization via declarative instructions on 1600+ tasks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.638882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.638882Z digest=sha256:2e5c99a87393d923e7367d25167c6999019724c978e995ff3bca71e0705aa4c3

Observation 86907fa7-da6a-4695-adfa-7f5c27498499 · outbound

This paper cites Disagreement Matters: Preserving Label Diversity by Jointly Modeling Item and Anno- tator Label Distributions with DisCo.

A Framework for Evaluating LLMs Under Task Indeterminacy Disagreement Matters: Preserving Label Diversity by Jointly Modeling Item and Anno- tator Label Distributions with DisCo

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.643852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.643852Z digest=sha256:5d4249a5518f36721d0710fcdde1f28430bf01625e9dba1893f5b478b9918a1e

Observation 35c8acda-de24-46ca-a0e7-f120e657f2f7 · outbound

This paper cites Many islands, many problems: An empirical examination of online safety behaviors in the caribbean.

A Framework for Evaluating LLMs Under Task Indeterminacy Many islands, many problems: An empirical examination of online safety behaviors in the caribbean

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.555782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:58:47.648449Z digest=sha256:dfd91e5354c29f268f4ecca244e8a4267a34d973fbb31a73ea4e1aad4fb0b33c

Pith citing papers

No inbound Pith citation observations are available.