Pith. sign in

Extracting Prompts by Inverting LLM Outputs

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We consider the problem of language model inversion: given outputs of a language model, we seek to extract the prompt that generated these outputs. We develop a new black-box method, output2prompt, that learns to extract prompts without access to the model's logits and without adversarial or jailbreaking queries. In contrast to previous work, output2prompt only needs outputs of normal user queries. To improve memory efficiency, output2prompt employs a new sparse encoding techique. We measure the efficacy of output2prompt on a variety of user and system prompts and demonstrate zero-shot transferability across different LLMs.

fields

cs.CR 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

SoK: Semantic Privacy in Large Language Models

cs.CR · 2025-06-30 · conditional · novelty 4.0

A systematization of knowledge arguing that LLM privacy threats extend beyond data leakage to semantically inferred attributes, and that current defenses only partially address them.

citing papers explorer

Showing 1 of 1 citing paper.

  • SoK: Semantic Privacy in Large Language Models cs.CR · 2025-06-30 · conditional · none · ref 47 · internal anchor

    A systematization of knowledge arguing that LLM privacy threats extend beyond data leakage to semantically inferred attributes, and that current defenses only partially address them.