Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:58:47.648449Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2411.13760.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:58:47.648449Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5c077c67-c093-4ebf-88d3-ecb29673ac03 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy DICES Dataset: Diversity in Conversational AI Evaluation for Safety
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a57fd76d-a59b-4874-9ec1-70b47786c546 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Stop Measuring Calibration When Humans Disagree
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b3da26a-e236-43ed-9e86-cac277deb601 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy It’s the End of the Gold Standard as we Know it
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8d37a0ce-5096-4610-a17a-d9fba6cdd5c2 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Like trainer, like bot? Inheritance of bias in algorithmic content moderation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 28c48695-e4b3-47dc-8a24-ac5d9e0ce8b6 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Holistic evaluation of language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 28671047-f62c-4e76-bf92-a379645b173b · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00d71071-c431-4e74-8027-1c47e9925021 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Revolt: Collaborative Crowdsourcing for Labeling Machine Learning Datasets
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6af6ccc-9894-468b-b70a-b43f38dc4034 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Judgment sieve: Reducing uncertainty in group judgments through interventions targeting ambiguity versus disagreement
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9a458d73-d10a-4790-80a3-159d996761a7 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Judgment Sieve: Reducing Uncertainty in Group Judgments through Interventions Targeting Ambiguity versus Disagreement
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9b8ed86a-925b-4376-959a-e23aace9ef7d · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Dealing with Disagree- ments: Looking Beyond the Majority V ote in Subjective Annotations
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ff854f86-582a-476c-b805-f332c7290525 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and Evaluation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ea58abb-d693-4056-aac1-58b585a5ad0a · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Red-Teaming for Generative AI: Silver Bullet or Security Theater?
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aed3702-9d7b-445e-aefd-781624d01b0f · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Efficient Conformal Prediction via Cascaded Inference with Expanded Admission
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ea16206-678b-402d-8ee8-379a7cd8505f · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb126527-a912-4f1a-a3e6-558a7fc321a7 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68160aa2-5850-4122-90bc-0f1ae0396ae7 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Deep Label Distribution Learning With Label Ambiguity
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4809e15c-34b0-4d78-b633-a8e1ed9dc364 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Label Distribution Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 210bc53a-7606-4252-9fbc-38252d34451d · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy The disagreement deconvolution: Bringing machine learning performance metrics in line with reality
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a7959e7b-177f-4f5d-a252-6e4a117e8c02 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7d6d6bd-d8b1-45c2-9a75-a31e0516dfad · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Jury learning: Integrating dissenting voices into machine learning models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9cbcbbe0-069c-473b-9568-f8cd5de417a1 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 50d372cd-4d44-4a5d-aa16-ca30d7b7bb7b · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Measuring Massive Multitask Language Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28b35d4a-d867-48bc-9e37-4ffe9a43f1d8 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Intersectionality in ai safety: Using multilevel models to understand diverse perceptions of safety in conversational ai
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1bf6244c-5633-479d-990b-713b16189533 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Culturally Aware Natural Language Inference
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77dd0d86-8374-4649-9493-e5226e7d39dd · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Chaithanya Manam, Dwarakanath Jampani, Mariam Zaim, Meng-Han Wu, and Alexander J
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a874b2e7-e4c3-487c-88ed-ec1d04b97ef8 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Annotation Error Detection: Analyzing the Past and Present for a More Coherent Future
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d383fbc1-768b-4684-90d8-75a83dca527e · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy A Bayesian Framework for Modeling Human Evaluations
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3ef8992-3d96-46b0-ad4e-d464487125d0 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Reconsidering Annotator Disagreement about Racist Language: Noise or Signal?
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c16a2fd9-fa2f-488f-aada-166b24f23f4d · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Learning to predict population-level label distributions
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f5c68c08-3d94-4613-8b4f-72dae82cb0c8 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy A Safe Harbor for AI Evaluation and Red Teaming
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b032cac8-7e8e-4f0f-ac4b-c349810d2c66 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy A Framework for Automated Measurement of Responsible AI Harms in Generative AI Applications
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9a420e2-695e-46aa-97db-9bd5713d93c7 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 630a4f9b-ddfe-4f13-bd99-180ec7bc36f4 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy StereoSet: Measuring stereotypical bias in pretrained language models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6657a046-e9a1-4318-a733-e6ca4ee0f37e · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Diversity-aware annotation for conversational ai safety
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a6e9d212-c5aa-486e-ac2b-6bb52305a3f3 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Inherent disagreements in human textual inferences
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4540f88d-e957-4faf-8bf6-61f13e4b7019 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Inherent Disagreements in Human Textual Inferences
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea6b370c-3186-416d-83f1-dc46594621fe · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Human uncertainty makes classification more robust
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c67e5e5-81c0-44dc-87a4-d46a433706d4 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy The 'Problem' of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 272987e6-d92a-44e6-a2a4-c316535e7c9a · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy On Releasing Annotator-Level Labels and Information in Datasets
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1e7394c-d3f0-43e8-b506-d192bc39d250 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Principled Evaluation with Human Labels: One Rater at a Time and Rater Equivalence
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 602c5eaf-4359-4f96-8eaa-e3f5f29ae945 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09088a7d-bbd7-4f5d-ac12-f8a906a710b3 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83efe431-4f75-4d3f-a57a-9139d956405b · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy A case for soft loss functions
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5c52f95-88a5-46ce-9d23-698238ed1636 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation affa9a7a-e8f2-410a-a2c9-bcd63e415c00 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy gold data
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 993a6c94-0d98-48a6-ac30-d1adeb6b00b7 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Super- naturalinstructions:generalization via declarative instructions on 1600+ tasks
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86907fa7-da6a-4695-adfa-7f5c27498499 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Disagreement Matters: Preserving Label Diversity by Jointly Modeling Item and Anno- tator Label Distributions with DisCo
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35c8acda-de24-46ca-a0e7-f120e657f2f7 · outbound
A Framework for Evaluating LLMs Under Task Indeterminacy Many islands, many problems: An empirical examination of online safety behaviors in the caribbean
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.