Pith. sign in

Paper Citation Record · LEDGER

Quantifying and Predicting Disagreement in Graded Human Ratings

As of 5 August 2026, this Paper Citation Record lists 100 of 220 outbound references and 1 inbound Pith citation observation for arXiv:2605.01168.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.01168 v1

Coverage vector

measured 100 of 220 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T18:35:58.177516Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T22:08:39.405870Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 220 outbound references displayed

  • verified exact20
  • verified fuzzy77
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4040fad5-9d13-4290-8cbe-1b7979585771 · outbound

This paper cites Attention is All you Need , url =.

Quantifying and Predicting Disagreement in Graded Human Ratings Attention is All you Need , url =

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.226579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:edfbf8cabfdbf63646837c7aa110f4dce805b3d24ee10f8af56850c4ca0ed3c4

Observation c71688ce-0a74-4770-a2b1-ba6da0e174a2 · outbound

This paper cites Seventeenth Symposium on Usable Privacy and Security (SOUPS 2021) , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Seventeenth Symposium on Usable Privacy and Security (SOUPS 2021) , pages=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.209942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:157be5496d5ad007927a9a20ce688a4142d34f6120f8055d8236f15f857cba86

Observation c99b66c2-f1b8-4cda-a55b-c55b54894354 · outbound

This paper cites Stop measuring calibration when humans disagree.

Quantifying and Predicting Disagreement in Graded Human Ratings Stop measuring calibration when humans disagree

Reference 3

Resolution
verified exact
doi, observed 2026-05-09T19:40:40.223394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:505106cb85be4b8573f7f3d20281255366c0cd55018deb04c37381af3af50573

Observation 7a7adb1e-426e-4e57-a555-6f6ceaacf944 · outbound

This paper cites Cognitive psychology , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Cognitive psychology , volume=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.206084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:7f7643baf778debb7b5ec139f26255e32467fb27c260e21311d2fe51fd01ef5c

Observation a8ef98cc-28b6-4ef9-8ecf-50bdb4f21128 · outbound

This paper cites Cognitive development and acquisition of language , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Cognitive development and acquisition of language , pages=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.213527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:df368c0cd1f9794ee71e36f10a7aa9b7700b1c26b6799db86683dc6591dea85b

Observation 230426ef-67ac-4e3c-ae53-c4beff17b302 · outbound

This paper cites Annual review of sociology , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Annual review of sociology , volume=

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.193176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:e77bd411d95b86be0bbd6caa1f9dc038b39c40a37f5c9410efe91c148f23a0cb

Observation 42df467b-e3c0-496b-a27e-5ee953c55782 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop) , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop) , pages=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.186096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:f33fc91fa6b29b6aa3dde716b056023fa9078a1a290d2d7e56e30ffae373a301

Observation f58554b5-1056-4729-8379-81546549b5a0 · outbound

This paper cites Modular pluralism: P luralistic alignment via multi- LLM collaboration.

Quantifying and Predicting Disagreement in Graded Human Ratings Modular pluralism: P luralistic alignment via multi- LLM collaboration

Reference 9

Resolution
verified exact
doi, observed 2026-05-09T19:40:40.215766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:fc2156cc2fb9f7f8b1297d7e514c2a92b4e1d5ea9970221a2a9c1c68922adfb7

Observation bac57fa7-6630-4ff3-b168-57852407d056 · outbound

This paper cites Mobile DNA , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Mobile DNA , volume=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.172693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:5c12f383fa606dad8940a2b85c890705ba477b7f8ddb9c617bd9bfe4bbd3b2ce

Observation 821bf2f4-6e03-4289-8db5-60d26274adc3 · outbound

This paper cites Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' Disagreement.

Quantifying and Predicting Disagreement in Graded Human Ratings Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' Disagreement

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.872113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:9cf9c25f239eef09bbcf643b18128eb0338ce93dd8f6e4570d639e7350ad2b65

Observation 67bb24c3-b836-45ee-a8b0-83bcbf89441a · outbound

This paper cites CEUR WORKSHOP PROCEEDINGS , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings CEUR WORKSHOP PROCEEDINGS , volume=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.164528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:8d98612b633c6412d59388288aa11e721d4e073a25363c7fc4cc84c78fa1f314

Observation 0d4a6f71-815c-4e2d-a4a3-f860f12908de · outbound

This paper cites an unresolved cited work.

Quantifying and Predicting Disagreement in Graded Human Ratings Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-25T23:01:20.180088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:c506f17f0b22aa56c95d4217a1dd7c4e9d722c51709a89052d6d8acb8fa579fe

Observation 89922f2b-8494-41a8-96fd-8344ee69ea28 · outbound

This paper cites European Conference on Information Retrieval , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings European Conference on Information Retrieval , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.202044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:a90a6e338c4fe236ea5ddc26fa451a09ce37c6d217497e7e22337c7104f3a59f

Observation e0876879-c7e4-49df-87b5-f9fdfa0c4260 · outbound

This paper cites Working Notes of CLEF , year=.

Quantifying and Predicting Disagreement in Graded Human Ratings Working Notes of CLEF , year=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.217604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:7296f3b34fd863fd7b8216e91f15815f7d8847f308be6220afa1cb2ea7036a07

Observation fef0a16a-067b-4d69-9b38-fbe5cf7a522d · outbound

This paper cites Proceedings of the 3rd Workshop on Perspectivist Approaches to NLP (NLPerspectives)@ LREC-COLING 2024 , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of the 3rd Workshop on Perspectivist Approaches to NLP (NLPerspectives)@ LREC-COLING 2024 , pages=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.235435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:fff24843f4a55be0f1a2784536436e7ae44acbec319abbe7f0fa8316a57915dd

Observation 50023a21-246b-4dce-93cc-d7dc97cb7cd6 · outbound

This paper cites Proceedings of the 9th ACM Multimedia Systems Conference , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of the 9th ACM Multimedia Systems Conference , pages=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.277857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:d8b4692834fa0d2b94c1f1837e3f8e2834c25c5306558377b2a09233b6ebed33

Observation d4728c83-3dd9-4343-9d9a-fec2686c342c · outbound

This paper cites IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , volume=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.358063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:ecc41d8e9b3eca3e08f583c9401bf79852cb5256a8921d8ea55ffe1c0cd8625b

Observation 1f0454a3-fe70-47d5-bb12-c53cf0566c81 · outbound

This paper cites In Search of Basic Units of Spoken Language , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings In Search of Basic Units of Spoken Language , pages=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.632168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:04d5eff3d78ecb60177e8905a649f12fef3dd02944850c59379ba66118c48d6f

Observation a4697f60-5fd1-49ed-8ca6-7a29ded4513c · outbound

This paper cites author=.

Quantifying and Predicting Disagreement in Graded Human Ratings author=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.696108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:1d9b32c3cb0dba4cdb57b8f7d709263ca5ae8b94b93416e50f7fe8af28f614ce

Observation 98a0a5a8-6697-4b47-ae85-cb5005fd2bb0 · outbound

This paper cites Neural Computing and Applications , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Neural Computing and Applications , volume=

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.639527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:37422b052e1d66c240ac48bc0606ebd2f6db39aa19c952cc7ec3ceab21aef1f9

Observation 0fdaf6f8-dfd5-4096-b444-fb27eb07d5cf · outbound

This paper cites Journal of New Music Research , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Journal of New Music Research , volume=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.720552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:faa096200aeceef33f5a034e16c31960f7321742555f2979d50271429a9a4ba9

Observation f67d1f7a-bdb0-49d8-ae75-aec78b0e28e0 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.731546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:14d86ddbb9a280b71ca4b018d959c76ea689fd1d6db1c583857eda476e7760c8

Observation c36e6b09-99ea-4cde-aa63-e8a6afcf5c9a · outbound

This paper cites an unresolved cited work.

Quantifying and Predicting Disagreement in Graded Human Ratings Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-25T23:01:20.197733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:ac758d53ef83a0f060efc70ee04414108ee849359cd09bd25801aac81fd0dc1e

Observation 212a94c2-64ec-4c7d-9fae-c57dcef48c15 · outbound

This paper cites Essentials of language documentation , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Essentials of language documentation , volume=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.739436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:9a4660619af544010a0e46fb2270f9fda5535c971f8ef15b43223afce62c4a5e

Observation 7323b5e2-2adc-40b0-844b-c60feec11fdd · outbound

This paper cites Computational linguistics , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Computational linguistics , volume=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.691340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:8ec0d6810ee9a400ae83f860b5fb49d65ae21c5c800922520e0672d66087e17d

Observation f71faaa9-235e-411b-aca2-576c1343db68 · outbound

This paper cites Finding Patterns in Noisy Crowds: Regression-based Annotation Aggregation for Crowdsourced Data.

Quantifying and Predicting Disagreement in Graded Human Ratings Finding Patterns in Noisy Crowds: Regression-based Annotation Aggregation for Crowdsourced Data

Reference 28

Resolution
verified exact
doi, observed 2026-05-09T19:40:40.213985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:562c2bdf9ff8cb8e929435a31125639489e77a47931346a21a10a63f3d85c662

Observation 902f49a8-fa09-425c-83ce-5f4f023b3400 · outbound

This paper cites Proceedings of The Web Conference 2020 , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of The Web Conference 2020 , pages=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.717616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:de99c11867660b12a13f117a1896f7c613697b979a9355ca7a2a75d29b785094

Observation 79eff9a3-0a08-4a2d-9908-10ce9e71b0e3 · outbound

This paper cites ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?.

Quantifying and Predicting Disagreement in Graded Human Ratings ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.840823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:ec779ab8b7e8ed17dd15ad3ffc1c1d5b3f67f73d1210fddfce94ddd5a3f7c96a

Observation c1ccacc8-fb01-4c11-ad25-4761f57ff9e9 · outbound

This paper cites 2009 IEEE conference on computer vision and pattern recognition , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings 2009 IEEE conference on computer vision and pattern recognition , pages=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.709624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:796d8ee1bf1a79925a54e5dab376bcd2c102cf3da18472eae105a1de58212762

Observation ecc82332-8535-4d81-99c7-1c0b6921db9a · outbound

This paper cites The state-of-the-art in.

Quantifying and Predicting Disagreement in Graded Human Ratings The state-of-the-art in

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.688925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:3eba7e492ae40b239eca28229c8200381b62caffea32f089058375f88854ff4f

Observation 3106dd56-4522-4f92-b052-bc16dc452b49 · outbound

This paper cites Procedia Computer Science , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Procedia Computer Science , volume=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.635825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:96f19283e6faa777ab071b665cb2dff474a5a8924fbadcae857d7e2903992f9b

Observation 98dc51f2-8ee1-42e9-9726-e1b4e21b3da6 · outbound

This paper cites Solving Label Variation in Scientific Information Extraction via Multi-Task Learning.

Quantifying and Predicting Disagreement in Graded Human Ratings Solving Label Variation in Scientific Information Extraction via Multi-Task Learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.643423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:2fc3a7e58a8105bf4a6c381dc150c4d50b4427ce000b082442b0f99a8049bebc

Observation d9eae6d9-b37c-4e07-ba96-a37b75614529 · outbound

This paper cites In Proceedings of the 3rd Workshop on Perspectivist Approaches to NLP (NLPerspectives)@ LREC-COLING 2024 , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings In Proceedings of the 3rd Workshop on Perspectivist Approaches to NLP (NLPerspectives)@ LREC-COLING 2024 , pages=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.685661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:5051fc21895c769261d5ee46c8907c650a3f2f1d27fe46292b1aec5df02efaea

Observation 4795be37-3ac0-489c-ac05-30747bf7cb75 · outbound

This paper cites Which Demographics do LLMs Default to During Annotation?.

Quantifying and Predicting Disagreement in Graded Human Ratings Which Demographics do LLMs Default to During Annotation?

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.993072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:49dd3ddc44df137243a617208efb1729a5144a99ff2497b5d67a5bd7821bf540

Observation b0b0edd9-47d0-4e9e-981e-c732861571e8 · outbound

This paper cites Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.691942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:78891afb440f77b689ecc578f7b5abbc22be58d1b947e7ccbfd7ad9284228070

Observation 79a7bfa4-4b6a-499b-b68f-69da167e298e · outbound

This paper cites Fine-grained Fallacy Detection with Human Label Variation.

Quantifying and Predicting Disagreement in Graded Human Ratings Fine-grained Fallacy Detection with Human Label Variation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.818082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:51f3c61972961b8df7cbbb1b955dfe6807e09e652a0696f8162dc4a94fe4c326

Observation d311c6e3-0f33-4056-a6b8-da97bdf74d83 · outbound

This paper cites Can Large Language Models Capture Dissenting Human Voices?.

Quantifying and Predicting Disagreement in Graded Human Ratings Can Large Language Models Capture Dissenting Human Voices?

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.757356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:add69ae5621cfbb1b8d0e3f00248d563265c7ededa735aa148ea667f6e6813c3

Observation edeb581c-9094-4cab-9722-f6b4fdb7f355 · outbound

This paper cites Proceedings of the Third Workshop on Understanding Implicit and Underspecified Language , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of the Third Workshop on Understanding Implicit and Underspecified Language , pages=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.687057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:293153fcc558d00b99b3a825da2d5e51ecfbc9dc85e2105645b17968c830b54e

Observation 061b7f92-624e-469a-862d-3645a983a948 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Transactions of the Association for Computational Linguistics , volume=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.315910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:e12a358daa4aa72a303d524add36246c7e683a035cc9308c0b836851b788a42a

Observation e0d85500-cfbf-4079-8299-f01b7bab4495 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Transactions of the Association for Computational Linguistics , volume=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.349967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:f2ba4dd41e109b73336f53d0325874762bb2f2e90a518da14ddbaabaecff3419

Observation a64d4595-b0b2-45b5-a8c8-750575ab7b0f · outbound

This paper cites Don't Blame the Data, Blame the Model: Understanding Noise and Bias When Learning from Subjective Annotations.

Quantifying and Predicting Disagreement in Graded Human Ratings Don't Blame the Data, Blame the Model: Understanding Noise and Bias When Learning from Subjective Annotations

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.984140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:28c89579befa2bcdfa903b93c73ecfd6b02acbbcea880df0f8a1b2b065b09de6

Observation 7069cacb-1931-4526-bf8b-3feabc0fc00d · outbound

This paper cites SemEval-2023 Task 11: Learning With Disagreements (LeWiDi).

Quantifying and Predicting Disagreement in Graded Human Ratings SemEval-2023 Task 11: Learning With Disagreements (LeWiDi)

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.792976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:bba190e1e031cc9499e60e22d1b7bda3be5926366f983ca5867c224aa06c31ac

Observation 2e29bcd9-1428-45a8-b6f0-8b0093e16676 · outbound

This paper cites Computational Linguistics , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Computational Linguistics , volume=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.452953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:828d9ca00d28c144286408b5cedb796e194d0ffeb5999e4c1c7c99152c9ad332

Observation e7f9dbf5-df03-416a-a609-1d4e0de78d60 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Transactions of the Association for Computational Linguistics , volume=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.506934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:f0b688b1cb6f150639b30ff225aa3dc967a69b436c8cadc81e3e65925066922c

Observation e98a6fb4-a83c-402a-9952-c29e45ae8a81 · outbound

This paper cites Computational Linguistics , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Computational Linguistics , volume=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.725333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:86a9c6a5f18d43856343b967243290784d595100cf1bb0d638d4185135c7845b

Observation dbc918ba-b997-462f-879e-4991c192d261 · outbound

This paper cites Language resources and evaluation , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Language resources and evaluation , volume=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.672620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:ed74d345b527f6aaee3941b643b33b5ace6ca62626cd536b3e729c413db9eda3

Observation 6f31983c-640c-4563-a82f-1fcce9cfab58 · outbound

This paper cites Learning with Annotation Noise.

Quantifying and Predicting Disagreement in Graded Human Ratings Learning with Annotation Noise

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.667728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:6309879d3147de51123648e27fe453f9d4dee28e86b2f5e66612a0a47ea9a6f4

Observation e32dd7e1-6b80-4763-8d35-396ac5f110a8 · outbound

This paper cites NLP ositionality: Characterizing Design Biases of Datasets and Models.

Quantifying and Predicting Disagreement in Graded Human Ratings NLP ositionality: Characterizing Design Biases of Datasets and Models

Reference 50

Resolution
verified exact
doi, observed 2026-05-09T19:40:40.217496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:448ad0b914d65e0b138d81a419885cef990a97088665b387374855daefe71f75

Observation 2bc9f035-de32-4c53-b4c5-a39842938868 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Transactions of the Association for Computational Linguistics , volume=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.655209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:62c141cd470fc82de5a085f4dce3e237451edb8ea668918c8eb2d58750f1307f

Observation ecec3f8c-7bd3-4535-a938-6af91fef85c7 · outbound

This paper cites L a MP : When Large Language Models Meet Personalization.

Quantifying and Predicting Disagreement in Graded Human Ratings L a MP : When Large Language Models Meet Personalization

Reference 52

Resolution
verified exact
doi, observed 2026-05-09T19:40:40.219119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:4e9bc3db27670a995f52f902d66932176840e3fe8080e20dfc178c8b4d7c96be

Observation fa3eca97-e4fd-46ff-94f2-68f1c7dc9424 · outbound

This paper cites Alfonso and Martin, Maite.

Quantifying and Predicting Disagreement in Graded Human Ratings Alfonso and Martin, Maite

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.649368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:b793c41b7b874697a262509d11a2add26a7d8d7e242dbbc7a6e9e6cbbb81cfeb

Observation 092257f9-d31c-4503-9ed1-f045cad86f76 · outbound

This paper cites arXiv preprint arXiv:2301.10684 , year=.

Quantifying and Predicting Disagreement in Graded Human Ratings arXiv preprint arXiv:2301.10684 , year=

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.765681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:3c4480027575aa85f23f2d870e5480d4ac7594d3dbc073db9fb2b659f02e5d52

Observation 05faa2b3-0722-4698-b143-901c2812c05a · outbound

This paper cites Disagreement in Argumentation Annotation.

Quantifying and Predicting Disagreement in Graded Human Ratings Disagreement in Argumentation Annotation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.727223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:b8195dca58ebd7949dcb97a8866fa35765d09c893058fc8072c553529b802fe1

Observation 02540f8b-aa42-4c5d-a677-5073e04e0518 · outbound

This paper cites When Do Annotator Demographics Matter? Measuring the Influence of Annotator Demographics with the POPQUORN Dataset.

Quantifying and Predicting Disagreement in Graded Human Ratings When Do Annotator Demographics Matter? Measuring the Influence of Annotator Demographics with the POPQUORN Dataset

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:09.006148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:ee4af55fea9de727f12c1e98ca25e7e2eae6415a959b6d003cae5d005a9838a8

Observation 711c9314-9f73-4406-9a8a-35a62522dfd3 · outbound

This paper cites Victoria and Herrera, Francisco.

Quantifying and Predicting Disagreement in Graded Human Ratings Victoria and Herrera, Francisco

Reference 57

Resolution
verified exact
doi, observed 2026-05-09T19:40:40.210425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:67ba48c149694e63ebe3b1b420a7151ccd4c75e5de67ff146050862fd1fb7306

Observation d51b3aed-5706-4a85-81ff-57822f98fe5d · outbound

This paper cites Proceedings of the 1st workshop on benchmarking: past, present and future , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of the 1st workshop on benchmarking: past, present and future , pages=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.735931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:a36150ccddce7a1d9248b5580e1bdd3eac65e1853d5c99fec02a2e9b541830b8

Observation 77b89ef4-b76d-425b-9e36-6131721332d9 · outbound

This paper cites Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting.

Quantifying and Predicting Disagreement in Graded Human Ratings Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.716677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:7c16c51edf2a974e37ff1cb5f778a35357a3483077b2d8fc1674859cf44899e9

Observation 758f0ab1-3ddd-4e7a-806a-634784ff45c9 · outbound

This paper cites Quantifying the persona effect in LLM simulations.

Quantifying and Predicting Disagreement in Graded Human Ratings Quantifying the persona effect in LLM simulations

Reference 60

Resolution
verified exact
doi, observed 2026-05-09T19:40:40.212155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:691d8056fef4a05541935c4e38c90b3be6a40c9c74a18d67631a18235626f65a

Observation 4b45e202-eca6-40cf-8d05-8b2da0a79aa5 · outbound

This paper cites Proceedings of the ACM on Human-Computer Interaction , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of the ACM on Human-Computer Interaction , volume=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.700654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:4752ca3789f9f9cfd6692d719da9c06bbead42b231c4882681be8fd473156835

Observation 10a0e9fa-2380-4369-a2cb-11d4301d860e · outbound

This paper cites PloS one , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings PloS one , volume=

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.713263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:f8f4b3375bdc00379ee638e4298e783bc244502a000ec2ecd4269faf83f9d3e5

Observation 20e93271-47e7-4aba-b27e-6179cb045f7b · outbound

This paper cites Journal of Communication Inquiry , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Journal of Communication Inquiry , volume=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.724144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:0e210b12932e963b4b47706f2328bbae0c49a1cb8a335dacdfe70635d4970d2e

Observation 230ebfa3-eeed-404b-9d14-187bfb479880 · outbound

This paper cites Proceedings of the Ninth International Workshop on Natural Language Processing for Social Media , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of the Ninth International Workshop on Natural Language Processing for Social Media , pages=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.721723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:1e713706cdb584cada6ba87ccb5155610773b98a378395e0d18600a196c30b6a

Observation 5a672512-ff5b-4a39-93c4-f510ebf27823 · outbound

This paper cites Annotating.

Quantifying and Predicting Disagreement in Graded Human Ratings Annotating

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.611498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:8f8ce1f7373e7f7a20377cb930dd9bd6f7d7a554a62aa0deabc5a4bac75707c3

Observation 3ffedcd1-f1f8-43a1-a502-ccf0d8aa44f5 · outbound

This paper cites Proceedings of the 7th International Conference on Natural Language and Speech Processing (ICNLSP 2024) , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of the 7th International Conference on Natural Language and Speech Processing (ICNLSP 2024) , pages=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.695242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:c3a92de9d3faa93106e894637faf87f2040c21a11fdc300d7a85fc9c667bce08

Observation d6bb5643-e38a-4daa-ac0e-9c911e71e739 · outbound

This paper cites Proceedings of the 15th ACM Web Science Conference 2023 , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of the 15th ACM Web Science Conference 2023 , pages=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.709241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:f37e9f8c574b20821f2142b613a9b92d95fce504c4f89fa6baa01c13b3c43e04

Observation b218f7b7-0812-4447-b54d-8d7e0294d400 · outbound

This paper cites 1st Workshop on Perspectivist Approaches to NLP , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings 1st Workshop on Perspectivist Approaches to NLP , pages=

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.666666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:9c5c24701984c4aaf0a4c0aec1023cb4c6af08f71318e7dfccfb71e0a889c309

Observation 62043f19-17fb-48bc-922b-c019423ff1f1 · outbound

This paper cites an unresolved cited work.

Quantifying and Predicting Disagreement in Graded Human Ratings Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-25T23:01:20.669731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:58b8a7ccc6bee62786ed8d1d95b236dbaef17694fcdfa6ff360f48f987b9f876

Observation 99e0877b-5691-4d0f-a93b-6c15f6cd0f89 · outbound

This paper cites Dialogue & Discourse , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Dialogue & Discourse , volume=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.675988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:4054e7a1fae39433526d0cd9c4a8a300e962f28272ffa5535dc67e0189bbad4f

Observation 3b380d06-56bf-4e93-9de5-cfd7c5f38883 · outbound

This paper cites Exploiting ` Subjective ' Annotations.

Quantifying and Predicting Disagreement in Graded Human Ratings Exploiting ` Subjective ' Annotations

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.678850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:054e07f3ecd554dcd5505878e44ba074ce73a199d42282a76908028d3825be8b

Observation 22a1ff07-f5c6-47fa-b02b-d3ff05e4616d · outbound

This paper cites Language Resources and Evaluation , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Language Resources and Evaluation , pages=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.652110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:f237b28e792740868653d39bbd452836154d1741b415938c2bcd18f117af8aab

Observation 6025bcd7-361b-4dc9-bf9c-5550a28e8615 · outbound

This paper cites The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels.

Quantifying and Predicting Disagreement in Graded Human Ratings The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.659551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:b9e44ff327af3d6547090599c63e6c5033201645d4cba59e59240a6e11d80754

Observation ffda838c-0053-4a40-a74f-4d6ecf3d4725 · outbound

This paper cites Proceedings of LAW X: The 10th Linguistic Annotation Workshop , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of LAW X: The 10th Linguistic Annotation Workshop , pages=

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.655821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:52b3df3207fcb8767b08598136cadbbdf2215f1511b6369a8c76252bfe8033c8

Observation 25ca8f1a-00f5-48fd-991f-8a9d6830080d · outbound

This paper cites Annotate.

Quantifying and Predicting Disagreement in Graded Human Ratings Annotate

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.698422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:17da6090e19d4cc2137778a01ffbe414a751cf03146cad16b7ea7c0c76df997d

Observation 5f7ef0d1-be94-44da-baa5-4443e78cf4e4 · outbound

This paper cites Language Resources and Evaluation , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Language Resources and Evaluation , volume=

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.682266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:ed672e927b90f495e0176c9ba1279090eb1a88f8f401f29528e748fa36c2ccc7

Observation 14bfe11c-d403-40a7-ae2b-a3004a33e3aa · outbound

This paper cites Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.701668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:b2dfe14248e03505013907e2bfeb77f83940668a9908eda8e2355edf0892959d

Observation 79d8e997-69e9-472f-998b-824f3ca66ea9 · outbound

This paper cites Different Tastes of Entities: Investigating Human Label Variation in Named Entity Annotations.

Quantifying and Predicting Disagreement in Graded Human Ratings Different Tastes of Entities: Investigating Human Label Variation in Named Entity Annotations

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.801213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:9761dee103699ffb7d96dacb9b9931d4fdb4d3f1bf9f407019f10740b5772ca0

Observation b5507d3d-6a4d-4efe-907e-b878ad33653b · outbound

This paper cites Architectural Sweet Spots for Modeling Human Label Variation by the Example of Argument Quality: It`s Best to Relate Perspectives!.

Quantifying and Predicting Disagreement in Graded Human Ratings Architectural Sweet Spots for Modeling Human Label Variation by the Example of Argument Quality: It`s Best to Relate Perspectives!

Reference 79

Resolution
verified exact
doi, observed 2026-05-09T19:40:40.208631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:c7b02c2c29e2ed807ccc5a92e3b5888b23e14ed902c8406e47d5501aa26ccd57

Observation 7c0b641b-ac4d-4599-bd88-12e1e8a6d5d4 · outbound

This paper cites Journal of Information Processing , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Journal of Information Processing , volume=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.623874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:fb97f8a81d449319ffe1ccc7dc377df1a0cca9b2ac92a3f5c7289938ab048853

Observation aa45c7ee-25af-4756-8b62-8154508f2eec · outbound

This paper cites Proceedings of the AAAI Conference on Human Computation and Crowdsourcing , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of the AAAI Conference on Human Computation and Crowdsourcing , volume=

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.610435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:8df1a675b6c8bd721a8034bba30eaf894678795c8317cd481f7a34d063d285a4

Observation ced4cf50-6975-461e-9573-5edbb534dcf7 · outbound

This paper cites The Importance of Modeling Social Factors of Language: Theory and Practice.

Quantifying and Predicting Disagreement in Graded Human Ratings The Importance of Modeling Social Factors of Language: Theory and Practice

Reference 82

Resolution
verified exact
doi, observed 2026-05-09T19:40:40.227412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:52ac4ce69f7b0113921f9dc44d8dce3dae93a2dc83ebf77ac36df0882d73d24e

Observation dabca256-e075-4fd2-90d8-7c82bf3b265b · outbound

This paper cites Identifying and Measuring Annotator Bias Based on Annotators ' Demographic Characteristics.

Quantifying and Predicting Disagreement in Graded Human Ratings Identifying and Measuring Annotator Bias Based on Annotators ' Demographic Characteristics

Reference 83

Resolution
verified exact
doi, observed 2026-05-09T19:40:40.221182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:d631920eb5db62607fefb2f2d88ad1ae6266c1a15023833b85dd2906e19e46f9

Observation fb7650da-d9e5-4a4d-bcdd-aa57ddcdef72 · outbound

This paper cites Information Processing & Management , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Information Processing & Management , volume=

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.614092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:dd3c77e6ea5aaf7117ec2c7e393045565e849dbc521d26b8ff789f5d0ea4585c

Observation d074b9c9-89dc-422b-9a20-39922129b0c3 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.601155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:e5f3a34c7f20e4ddca1be538796def48c2882ab2bb18880b248b1500ec7ec003

Observation 9778caef-3bf8-4298-a2ff-7705a73b7ae6 · outbound

This paper cites Computational Linguistics , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Computational Linguistics , volume=

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.597313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:dc4fb607a6304038963fb9a920fc86c62404e8f0fa623ce568b78dcd96c2ca29

Observation 1f73d892-8e37-4d53-8647-680a022e1f7b · outbound

This paper cites Subjective Natural Language Problems: Motivations, Applications, Characterizations, and Implications.

Quantifying and Predicting Disagreement in Graded Human Ratings Subjective Natural Language Problems: Motivations, Applications, Characterizations, and Implications

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.605571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:1af2a01efb8e0f3570b88c57b8a5c0e98286a8908fe67df4ccb33e1508ba7b9f

Observation 500be185-81be-40dc-afaf-373af35b5cdf · outbound

This paper cites Artificial intelligence and law , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Artificial intelligence and law , volume=

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.617570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:443934c7a2148fdab8ed6cad0f198a6fced277dda3b459f5000564a70ed88185

Observation 74c5a8e1-7faa-4d2f-a977-ca3124a18a87 · outbound

This paper cites Applied Clinical Informatics , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Applied Clinical Informatics , volume=

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.620740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:a18585ea3458484ab0a5b227403ba0c694503239a56a24ea1e1a5fb4372118ba

Observation 41684f7e-bd17-4e7f-9243-04cc63fa1677 · outbound

This paper cites author=.

Quantifying and Predicting Disagreement in Graded Human Ratings author=

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.639674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:c12d3490d276faf8351aab9a2ccffccbbf2bfdafd4011ad1be991b620468e4aa

Observation 1cbbffb5-055d-4bcb-9669-588792ddc20d · outbound

This paper cites Plos one , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Plos one , volume=

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.705334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:1f9c81e30311b95fc3fda744f53655060f4bd470b5ab6e6eeaf63ec2e8ecca62

Observation c8cb3227-08f1-43cb-a78d-daf48614a241 · outbound

This paper cites Computational linguistics , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Computational linguistics , volume=

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.581380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:79cfa8c8e1257fcdf7541495b9af739668365b0dcbc9361eb2262ac4c267efbd

Observation c6198c39-13fa-4d46-a664-872daffb8f3d · outbound

This paper cites Language resources and evaluation , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings Language resources and evaluation , volume=

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.589425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:9b3018bdbe8babfd0e45ea7c1a2f3339b9a12dfb403ba2cbffc1ad065eaee988

Observation 0ced1c8c-c6ad-4e30-acf2-2ae91c07142f · outbound

This paper cites Why Don ' t You Do It Right? Analysing Annotators ' Disagreement in Subjective Tasks.

Quantifying and Predicting Disagreement in Graded Human Ratings Why Don ' t You Do It Right? Analysing Annotators ' Disagreement in Subjective Tasks

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.585740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:3a2971ebaf6a0248ff30ce0d11c1f156b4b3e21046d3d05c30a1ca0f013122f7

Observation 29ec9b8e-1863-4a73-81ae-5e9e6d49e926 · outbound

This paper cites AI Magazine , volume=.

Quantifying and Predicting Disagreement in Graded Human Ratings AI Magazine , volume=

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.648579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:80570b645711de3635dcb60f4c977cb3a25d72daa5960bae1671fd30fe9eab8f

Observation f164d148-04f3-40dd-bc64-4535d4606257 · outbound

This paper cites The Semantic Web.

Quantifying and Predicting Disagreement in Graded Human Ratings The Semantic Web

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.682956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:3dcd3fd66cd3028720c388fa2b0694bb4ad8f9ecea8c05f943bd2f6d8f504826

Observation de3e7bea-2103-48d0-9231-780c43fe5a50 · outbound

This paper cites Investigating Reasons for Disagreement in Natural Language Inference.

Quantifying and Predicting Disagreement in Graded Human Ratings Investigating Reasons for Disagreement in Natural Language Inference

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:01:20.397550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:f666ad4f92215d27d8a068d5184adefb906e9a598af02a43f08fa765f3585190

Observation d61525ff-dff1-4832-949f-b3947ba96325 · outbound

This paper cites 2023 , booktitle=.

Quantifying and Predicting Disagreement in Graded Human Ratings 2023 , booktitle=

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.713967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:0258bea50733f27d6c55262965eea10379b6baee78b35372d1cabd8f0bcdfec5

Observation 723106d1-76f5-4304-9b18-a72aab84a15d · outbound

This paper cites Would you describe a leopard as yellow? Evaluating crowd-annotations with justified and informative disagreement.

Quantifying and Predicting Disagreement in Graded Human Ratings Would you describe a leopard as yellow? Evaluating crowd-annotations with justified and informative disagreement

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.678618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:643b0d08d1f79ff0ed0f81cbb125f2cb72da7ce9bfebf70f5040971b5006d128

Observation 89da5b02-f38d-4ca3-a324-c27423569f59 · outbound

This paper cites Embracing Ambiguity: A Comparison of Annotation Methodologies for Crowdsourcing Word Sense Labels.

Quantifying and Predicting Disagreement in Graded Human Ratings Embracing Ambiguity: A Comparison of Annotation Methodologies for Crowdsourcing Word Sense Labels

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.705291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:1e44faf1f830fcd12730658fb21c47f86ee3fbc02a4ff4eeab8e075de18ae8f7

Observation 77d0a823-7b9f-4069-96bd-2eded1c8e300 · outbound

This paper cites Annotating.

Quantifying and Predicting Disagreement in Graded Human Ratings Annotating

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.644256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:f812d4b7d2f93ae968974a75bd7c4dddd2192e1c896d10fc6750209b368dc746

Observation 2442e8d6-e8f5-447c-87ec-704f91c5f149 · outbound

This paper cites Are You a Racist or Am I Seeing Things? Annotator Influence on Hate Speech Detection on T witter.

Quantifying and Predicting Disagreement in Graded Human Ratings Are You a Racist or Am I Seeing Things? Annotator Influence on Hate Speech Detection on T witter

Reference 103

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T23:06:19.729383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:35:58.177516Z digest=sha256:01d4804fc3a8d2c13a7a1b80a12c82bc28a15bf2a63d5edec02c429dc17e43c2

Pith citing papers

Observation 93c846f6-8c33-4485-b814-7638ba7de141 · inbound

How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI cites this paper.

How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI Quantifying and Predicting Disagreement in Graded Human Ratings

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T22:08:39.405870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:08:39.405870Z digest=sha256:bdb6c406c78647d2086c4771868520e2f3ee485c63bc21ccb092357f040a423d