Pith. sign in

REVIEW 2 cited by

Equality before the Law: Legal Judgment Consistency Analysis for Fairness

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.13868 v1 pith:G2HFCALS submitted 2021-03-25 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords inconsistencyjudgmentlegaldatalincoconsistencydifferentgender
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In a legal system, judgment consistency is regarded as one of the most important manifestations of fairness. However, due to the complexity of factual elements that impact sentencing in real-world scenarios, few works have been done on quantitatively measuring judgment consistency towards real-world data. In this paper, we propose an evaluation metric for judgment inconsistency, Legal Inconsistency Coefficient (LInCo), which aims to evaluate inconsistency between data groups divided by specific features (e.g., gender, region, race). We propose to simulate judges from different groups with legal judgment prediction (LJP) models and measure the judicial inconsistency with the disagreement of the judgment results given by LJP models trained on different groups. Experimental results on the synthetic data verify the effectiveness of LInCo. We further employ LInCo to explore the inconsistency in real cases and come to the following observations: (1) Both regional and gender inconsistency exist in the legal system, but gender inconsistency is much less than regional inconsistency; (2) The level of regional inconsistency varies little across different time periods; (3) In general, judicial inconsistency is negatively correlated with the severity of the criminal charges. Besides, we use LInCo to evaluate the performance of several de-bias methods, such as adversarial learning, and find that these mechanisms can effectively help LJP models to avoid suffering from data bias.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    On 59k Canadian refugee decisions, three ML fairness approaches produce divergent, sometimes contradictory signals and fail to capture substantive legal reasoning.

  2. The Judge Variable: Challenging Judge-Agnostic Legal Judgment Prediction

    cs.CL 2025-07 reject novelty 4.0 of 10

    Models trained on individual judges' past child-custody rulings predict those judges' future rulings better than a model trained on all judges together, a result the paper reads as support for legal realism.

Pith tools