Pith. sign in

REVIEW 15 cited by

REEF: Representation Encoding Fingerprints for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.14273 v1 pith:UBO46EKP submitted 2024-10-18 cs.CL cs.AIcs.CR

classification cs.CLcs.AIcs.CR
keywords modelreefllmsmodelssuspectvictimidentifyintellectual
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Protecting the intellectual property of open-source Large Language Models (LLMs) is very important, because training LLMs costs extensive computational resources and data. Therefore, model owners and third parties need to identify whether a suspect model is a subsequent development of the victim model. To this end, we propose a training-free REEF to identify the relationship between the suspect and victim models from the perspective of LLMs' feature representations. Specifically, REEF computes and compares the centered kernel alignment similarity between the representations of a suspect model and a victim model on the same samples. This training-free REEF does not impair the model's general capabilities and is robust to sequential fine-tuning, pruning, model merging, and permutations. In this way, REEF provides a simple and effective way for third parties and models' owners to protect LLMs' intellectual property together. The code is available at https://github.com/tmylla/REEF.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. modelDNA: Calibrated Lineage Verification and Merge Decomposition from Sampled Weight Fingerprints

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Sampled weight fingerprints recover LLM parentage with AUROC 1.0 and zero false positives, and recover published mergekit mixture weights without full downloads.

  2. CTCC: A Robust and Stealthy Fingerprinting Framework for Large Language Models via Cross-Turn Contextual Correlation Backdoor

    cs.CL 2025-09 conditional novelty 6.0 of 10

    CTCC embeds LLM ownership fingerprints in cross-turn semantic contradictions: the model fires a secret response only when a user contradicts an earlier statement, with higher robustness and stealth than single-turn triggers.

  3. EverTracer: Hunting Stolen Large Language Models via Stealthy and Robust Probabilistic Fingerprint

    cs.CR 2025-09 conditional novelty 6.0 of 10

    A gray-box LLM fingerprinting method that detects memorized private text via calibrated probability variation, instead of using backdoor triggers.

  4. FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing

    cs.CR 2025-08 conditional novelty 6.0 of 10

    FPEdit uses knowledge editing with a promote-suppress objective to embed robust, stealthy natural-language fingerprints into LLMs, achieving 94 to 100 percent retention after fine-tuning while preserving benchmark per...

  5. Hot-Swap MarkBoard: An Efficient Black-box Watermarking Approach for Large-scale Model Distribution

    cs.CR 2025-07 conditional novelty 6.0 of 10

    A branch-swapping mechanism over low-rank add-on modules stamps each distributed model copy with a unique binary user ID, achieving 100% reported verification accuracy with under 1% extra parameters.

  6. Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space

    cs.AI 2026-08 conditional novelty 5.0 of 10

    Weight-space spectral statistics and subspace geometry separate independent, same-family, and shared-base LLMs, and track fine-grained post-training differences.

  7. MEraser: An Effective Fingerprint Erasure Approach for Large Language Models

    cs.CR 2025-06 conditional novelty 5.0 of 10

    By fine-tuning on mismatched pairs and then clean pairs, MEraser drops fingerprint success rate to zero on three backdoor-based fingerprinting schemes across multiple LLMs, with a reusable LoRA adapter for transfer.

  8. Gradient-Based Model Fingerprinting for LLM Similarity Detection and Family Classification

    cs.LG 2025-06 conditional novelty 5.0 of 10

    TensorGuard classifies fine-tuned LLMs into their base-model families with 94% accuracy by clustering statistical features of weight gradients under random input perturbations.

  9. Towards the Resistance of Neural Network Watermarking to Fine-tuning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    The authors prove that filter frequency components at certain frequencies are nearly unchanged by gradient descent when the layer input is low-frequency, and use these components as a fine-tuning-robust watermark.

  10. Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

    cs.CL 2025-02 conditional novelty 5.0 of 10

    TELLME edits an LLM's hidden representations so similar behaviors cluster and different behaviors separate, improving safety monitoring and detoxification while preserving general ability.

  11. I'm Spartacus, No, I'm Spartacus: Measuring and Understanding LLM Identity Confusion

    cs.CR 2024-11 reject novelty 5.0 of 10

    Seven of 27 tested LLMs (25.93%) exhibited identity confusion, which the authors link to hallucination and show reduces user trust, especially in critical tasks.

  12. Invisible Traces: Using Hybrid Fingerprinting to identify underlying LLMs in GenAI Apps

    cs.LG 2025-01 conditional novelty 4.0 of 10

    A hybrid of active and passive fingerprinting identifies the underlying LLM in simulated GenAI apps, reaching about 86.5% accuracy with ten observed responses.

  13. Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI

    cs.CY 2024-12 conditional novelty 4.0 of 10

    The paper proposes the AI-45 degree law, a Causal Ladder framework, and five trustworthiness levels as a roadmap toward trustworthy AGI.

  14. Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A survey maps the field of MLLM explainability and interpretability into data, model, and training and inference perspectives.

  15. Unlocking the Effectiveness of LoRA-FP for Seamless Transfer Implantation of Fingerprints in Downstream Models

    cs.CR 2025-08 conditional novelty 3.0 of 10

    Backdoor fingerprints trained into LoRA adapters on a base LLM transfer to derivative models with 100% trigger success and, in several scenarios, greater robustness than directly injected fingerprints.

Pith tools