Pith. sign in

REVIEW 1 cited by

Word Importance Explains How Prompts Affect Language Model Outputs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.03028 v1 pith:CINI5I4U submitted 2024-03-05 cs.AI cs.CL

classification cs.AIcs.CL
keywords importanceworddifferentimpactlanguagemultipleoutputsprompts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The emergence of large language models (LLMs) has revolutionized numerous applications across industries. However, their "black box" nature often hinders the understanding of how they make specific decisions, raising concerns about their transparency, reliability, and ethical use. This study presents a method to improve the explainability of LLMs by varying individual words in prompts to uncover their statistical impact on the model outputs. This approach, inspired by permutation importance for tabular data, masks each word in the system prompt and evaluates its effect on the outputs based on the available text scores aggregated over multiple user inputs. Unlike classical attention, word importance measures the impact of prompt words on arbitrarily-defined text scores, which enables decomposing the importance of words into the specific measures of interest--including bias, reading level, verbosity, etc. This procedure also enables measuring impact when attention weights are not available. To test the fidelity of this approach, we explore the effect of adding different suffixes to multiple different system prompts and comparing subsequent generations with different large language models. Results show that word importance scores are closely related to the expected suffix importances for multiple scoring functions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Meaningless is better: hashing bias-inducing words in LLM prompts improves performance in logical reasoning and statistical learning

    cs.CL 2024-11 conditional novelty 5.0 of 10

    Masking bias-triggering words with random identifiers increased accuracy on two small LLM reasoning and counting tasks, with effects varying by model.

Pith tools