Pith. sign in

REVIEW 1 cited by

Surfacing Biases in Large Language Models using Contrastive Input Decoding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.07378 v1 pith:473HMJOT submitted 2023-05-12 cs.CL cs.CYcs.LG

classification cs.CLcs.CYcs.LG
keywords decodinginputcontrastivegiveninputstextbiasesdifferences
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Ensuring that large language models (LMs) are fair, robust and useful requires an understanding of how different modifications to their inputs impact the model's behaviour. In the context of open-text generation tasks, however, such an evaluation is not trivial. For example, when introducing a model with an input text and a perturbed, "contrastive" version of it, meaningful differences in the next-token predictions may not be revealed with standard decoding strategies. With this motivation in mind, we propose Contrastive Input Decoding (CID): a decoding algorithm to generate text given two inputs, where the generated text is likely given one input but unlikely given the other. In this way, the contrastive generations can highlight potentially subtle differences in how the LM output differs for the two inputs in a simple and interpretable manner. We use CID to highlight context-specific biases that are hard to detect with standard decoding strategies and quantify the effect of different input perturbations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Amateur Contrastive Decoding for Text Generation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Multi-amateur contrastive decoding pools signals from a set of small language models, with mean or consensus aggregation, to improve open-ended text generation over single-amateur CD.

Pith tools