Pith. sign in

REVIEW 19 cited by

Stealing Part of a Production Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.06634 v2 pith:63EZNDZX submitted 2024-03-11 cs.CR

Stealing Part of a Production Language Model

classification cs.CR
keywords attacklanguagemodelmodelsprojectionblack-boxdimensionentire
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce the first model-stealing attack that extracts precise, nontrivial information from black-box production language models like OpenAI's ChatGPT or Google's PaLM-2. Specifically, our attack recovers the embedding projection layer (up to symmetries) of a transformer model, given typical API access. For under \$20 USD, our attack extracts the entire projection matrix of OpenAI's Ada and Babbage language models. We thereby confirm, for the first time, that these black-box models have a hidden dimension of 1024 and 2048, respectively. We also recover the exact hidden dimension size of the gpt-3.5-turbo model, and estimate it would cost under $2,000 in queries to recover the entire projection matrix. We conclude with potential defenses and mitigations, and discuss the implications of possible future work that could extend our attack.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Channel Location Constrains the Auditability of Subliminal Learning

    cs.LG 2026-06 unverdicted novelty 7.0

    Auditability of subliminal learning is constrained by channel location, with initialization-dependent body channels allowing pre-training screens while vocabulary geometry and conditional body channels evade them.

  2. SoK: Colluding Adversaries in Machine Learning Pipelines

    cs.CR 2026-06 unverdicted novelty 7.0

    The paper introduces a framework for collusion between train- and inference-time adversaries in ML pipelines, proposes a guideline for conjecturing collusion potential, explains prior work, and empirically validates f...

  3. Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data

    cs.LG 2026-05 unverdicted novelty 7.0

    ALU uses public data to suppress unlearning cost quadratically while characterizing distribution mismatch effects, enabling mass unlearning with maintained utility.

  4. Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data

    cs.LG 2026-05 unverdicted novelty 7.0

    Asymmetric Langevin Unlearning uses public data to suppress unlearning noise costs by O(1/n_pub²), enabling practical mass unlearning with preserved utility under distribution mismatch.

  5. A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework

    cs.CR 2026-04 unverdicted novelty 7.0

    A new 7x4 taxonomy organizes agentic AI security threats by architectural layer and persistence timescale, revealing under-explored upper layers and missing defenses after surveying 116 papers.

  6. Fingerprinting LLMs via Prompt Injection

    cs.CR 2025-09 conditional novelty 7.0

    LLMPrint generates unique, post-processing-robust fingerprints for base LLMs and their variants via optimized prompt injection with statistical verification for gray-box and black-box settings.

  7. Can Watermarking Techniques Help Prevent LLM Model Stealing?

    cs.CR 2026-07 conditional novelty 6.5

    Softplus-then-perturb with embedding-seeded Gaussian noise defeats PCA/averaging/RPCA dimension-extraction attacks on Mistral-7B and GPT-2 with only modest quality loss.

  8. Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions

    cs.AI 2026-07 conditional novelty 6.0

    Across Pythia models, a large fraction of training contexts are predicted almost exactly by the empirical next-token distribution of the corpus, but frequent high-entropy contexts remain poorly matched.

  9. Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions

    cs.AI 2026-07 conditional novelty 6.0

    A trained transformer's next-token distribution often matches the empirical next-token distribution of its pretraining corpus, with agreement improving as models grow, while a persistent tail of mismatches remains.

  10. Black-Box Inference of LLM Architectural Properties with Restrictive API Access

    cs.LG 2026-07 unverdicted novelty 6.0

    NightVision recovers LLM hidden dimension to 23% average relative error (9% on MoE) and depth/parameter count to 53% on models >3B parameters using common-set prompting, spectral analysis, and TTFT under single-logit ...

  11. Surrogate Fidelity: When Can Open LLMs Explain Closed Ones?

    cs.LG 2026-06 unverdicted novelty 6.0

    Prediction agreement between open and closed LLMs substantially overstates agreement on attributions and causal reasons.

  12. OTRO: Oblivious Tokenization Path with Square-Root ORAM

    cs.CR 2026-06 unverdicted novelty 6.0

    OTRO combines replicated square-root ORAM instances, epoch rotation with dummy padding, and KV-cache-aware chunking to make tokenizer lookups oblivious with at most 4.5% TTFT overhead and under 0.5 GB extra memory in ...

  13. The Surface You Test Is Not the Surface That Breaks

    cs.CR 2026-05 unverdicted novelty 6.0

    Prompt injection vulnerability in tool-augmented LLMs is a model-surface interaction rather than a fixed channel property; the same payload inverts success rates across models, and adaptive attack rate exceeds single-...

  14. On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference

    cs.CR 2026-05 conditional novelty 6.0

    An attack aligns differently shuffled intermediate activations from secure Transformer inference queries to recover model weights with low error using roughly one dollar of queries.

  15. Characterizing Linear Alignment Across Language Models

    cs.AI 2026-03 conditional novelty 6.0

    Linear (affine) maps between final hidden states of independent LLMs preserve downstream performance and can enable text generation when tokenizers and scale align, enabling a practical HE-based privacy protocol.

  16. CrypTorch: PyTorch-based Auto-tuning Compiler for Machine Learning with Multi-party Computation

    cs.CR 2025-11 conditional novelty 6.0

    An MPC-ML compiler that modularizes and auto-tunes operator approximations, delivering 1.2–1.8x speedups over an optimized baseline under user-set accuracy bounds.

  17. LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users

    cs.CL 2025-07 unverdicted novelty 6.0

    A single attacker can use strategic upvoting and downvoting on language model outputs to inject facts, security flaws, or fake news that persist in the model for all users after preference tuning.

  18. How Well Do AI Systems Solve AP Physics? A Comparative Evaluation of Large Language Models on Algebra-Based Free Response Questions

    physics.ed-ph 2026-03 unverdicted novelty 5.0

    ChatGPT 4.1 mini, Gemini 2.5 Flash, Claude 4.0 Sonnet, and DeepSeek R1 average 82–92% on AP Physics 1/2 free-response questions but systematically fail spatial, visual, and conceptual tasks.

  19. The Geometry of Last-Layer Model Stealing

    cs.LG 2026-06 unverdicted novelty 3.0

    Geometry maps the conditions for perfect last-layer theft in transformers and demonstrates that full hidden-network reverse engineering is impossible from final outputs.