Pith. sign in

REVIEW 1 cited by

Mitigating Cognitive Biases in Multi-Criteria Crowd Assessment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.18938 v1 pith:PBO3JL3D submitted 2024-07-10 cs.HC cs.LG

classification cs.HCcs.LG
keywords biasescognitiveaggregationassessmentcrowdsourcingcriteriaevaluationinter-criteria
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Crowdsourcing is an easy, cheap, and fast way to perform large scale quality assessment; however, human judgments are often influenced by cognitive biases, which lowers their credibility. In this study, we focus on cognitive biases associated with a multi-criteria assessment in crowdsourcing; crowdworkers who rate targets with multiple different criteria simultaneously may provide biased responses due to prominence of some criteria or global impressions of the evaluation targets. To identify and mitigate such biases, we first create evaluation datasets using crowdsourcing and investigate the effect of inter-criteria cognitive biases on crowdworker responses. Then, we propose two specific model structures for Bayesian opinion aggregation models that consider inter-criteria relations. Our experiments show that incorporating our proposed structures into the aggregation model is effective to reduce the cognitive biases and help obtain more accurate aggregation results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees

    cs.LG 2026-08 reject novelty 5.0 of 10

    Aggregate-then-Calibrate projects model scores onto a human-derived consensus ranking, claiming theoretical guarantees over model-only assessment; the central proofs have important gaps.

Pith tools