REVIEW 15 cited by
REEF: Representation Encoding Fingerprints for Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Protecting the intellectual property of open-source Large Language Models (LLMs) is very important, because training LLMs costs extensive computational resources and data. Therefore, model owners and third parties need to identify whether a suspect model is a subsequent development of the victim model. To this end, we propose a training-free REEF to identify the relationship between the suspect and victim models from the perspective of LLMs' feature representations. Specifically, REEF computes and compares the centered kernel alignment similarity between the representations of a suspect model and a victim model on the same samples. This training-free REEF does not impair the model's general capabilities and is robust to sequential fine-tuning, pruning, model merging, and permutations. In this way, REEF provides a simple and effective way for third parties and models' owners to protect LLMs' intellectual property together. The code is available at https://github.com/tmylla/REEF.
Forward citations
Cited by 15 Pith papers
-
modelDNA: Calibrated Lineage Verification and Merge Decomposition from Sampled Weight Fingerprints
Sampled weight fingerprints recover LLM parentage with AUROC 1.0 and zero false positives, and recover published mergekit mixture weights without full downloads.
-
CTCC: A Robust and Stealthy Fingerprinting Framework for Large Language Models via Cross-Turn Contextual Correlation Backdoor
CTCC embeds LLM ownership fingerprints in cross-turn semantic contradictions: the model fires a secret response only when a user contradicts an earlier statement, with higher robustness and stealth than single-turn triggers.
-
EverTracer: Hunting Stolen Large Language Models via Stealthy and Robust Probabilistic Fingerprint
A gray-box LLM fingerprinting method that detects memorized private text via calibrated probability variation, instead of using backdoor triggers.
-
FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
FPEdit uses knowledge editing with a promote-suppress objective to embed robust, stealthy natural-language fingerprints into LLMs, achieving 94 to 100 percent retention after fine-tuning while preserving benchmark per...
-
Hot-Swap MarkBoard: An Efficient Black-box Watermarking Approach for Large-scale Model Distribution
A branch-swapping mechanism over low-rank add-on modules stamps each distributed model copy with a unique binary user ID, achieving 100% reported verification accuracy with under 1% extra parameters.
-
Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space
Weight-space spectral statistics and subspace geometry separate independent, same-family, and shared-base LLMs, and track fine-grained post-training differences.
-
MEraser: An Effective Fingerprint Erasure Approach for Large Language Models
By fine-tuning on mismatched pairs and then clean pairs, MEraser drops fingerprint success rate to zero on three backdoor-based fingerprinting schemes across multiple LLMs, with a reusable LoRA adapter for transfer.
-
Gradient-Based Model Fingerprinting for LLM Similarity Detection and Family Classification
TensorGuard classifies fine-tuned LLMs into their base-model families with 94% accuracy by clustering statistical features of weight gradients under random input perturbations.
-
Towards the Resistance of Neural Network Watermarking to Fine-tuning
The authors prove that filter frequency components at certain frequencies are nearly unchanged by gradient descent when the layer input is low-frequency, and use these components as a fine-tuning-robust watermark.
-
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring
TELLME edits an LLM's hidden representations so similar behaviors cluster and different behaviors separate, improving safety monitoring and detoxification while preserving general ability.
-
I'm Spartacus, No, I'm Spartacus: Measuring and Understanding LLM Identity Confusion
Seven of 27 tested LLMs (25.93%) exhibited identity confusion, which the authors link to hallucination and show reduces user trust, especially in critical tasks.
-
Invisible Traces: Using Hybrid Fingerprinting to identify underlying LLMs in GenAI Apps
A hybrid of active and passive fingerprinting identifies the underlying LLM in simulated GenAI apps, reaching about 86.5% accuracy with ten observed responses.
-
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
The paper proposes the AI-45 degree law, a Causal Ladder framework, and five trustworthiness levels as a roadmap toward trustworthy AGI.
-
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey
A survey maps the field of MLLM explainability and interpretability into data, model, and training and inference perspectives.
-
Unlocking the Effectiveness of LoRA-FP for Seamless Transfer Implantation of Fingerprints in Downstream Models
Backdoor fingerprints trained into LoRA adapters on a base LLM transfer to derivative models with 100% trigger success and, in several scenarios, greater robustness than directly injected fingerprints.
Discussion (0). Continue with ORCID to comment.