Pith. sign in

REVIEW 3 cited by

Measuring the Inconsistency of Large Language Models in Preferential Ranking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08851 v1 pith:PVS4IBJU submitted 2024-10-11 cs.CL

classification cs.CL
keywords llmspreferentialalternativesbiasconsistentcriteriainconsistencyirrelevant
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite large language models' (LLMs) recent advancements, their bias and hallucination issues persist, and their ability to offer consistent preferential rankings remains underexplored. This study investigates the capacity of LLMs to provide consistent ordinal preferences, a crucial aspect in scenarios with dense decision space or lacking absolute answers. We introduce a formalization of consistency based on order theory, outlining criteria such as transitivity, asymmetry, reversibility, and independence from irrelevant alternatives. Our diagnostic experiments on selected state-of-the-art LLMs reveal their inability to meet these criteria, indicating a strong positional bias and poor transitivity, with preferences easily swayed by irrelevant alternatives. These findings highlight a significant inconsistency in LLM-generated preferential rankings, underscoring the need for further research to address these limitations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models

    cs.AI 2025-06 unverdicted novelty 6.0 of 10

    LLMs show a quality-dependent position bias, favoring the first option for high-quality choices and later options for low-quality ones, and higher-temperature sampling can reveal the underlying preference.

  2. Quantitative Relaxations of Arrow's Axioms

    cs.GT 2025-06 conditional novelty 6.0 of 10

    A new quantitative framework measures the degree to which voting rules violate Arrow's independence and unanimity axioms, and an empirical study finds Borda performs best on Scottish and synthetic elections.

  3. Agent Identity Evals: Measuring Agentic Identity

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Introduces Agent Identity Evals (AIE), five similarity-based metrics for LMA identity stability, with pilot experiments showing identifiability always at zero and no statistical support.

Pith tools