Pith. sign in

REVIEW 1 cited by

Low-rank Adaptation Method for Wav2vec2-based Fake Audio Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.05617 v1 pith:62MWWY3X submitted 2023-06-09 cs.SD cs.CLeess.AS

classification cs.SDcs.CLeess.AS
keywords fine-tuningmodelmodelsparameterspre-trainedtrainableadaptationaudio
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Self-supervised speech models are a rapidly developing research topic in fake audio detection. Many pre-trained models can serve as feature extractors, learning richer and higher-level speech features. However,when fine-tuning pre-trained models, there is often a challenge of excessively long training times and high memory consumption, and complete fine-tuning is also very expensive. To alleviate this problem, we apply low-rank adaptation(LoRA) to the wav2vec2 model, freezing the pre-trained model weights and injecting a trainable rank-decomposition matrix into each layer of the transformer architecture, greatly reducing the number of trainable parameters for downstream tasks. Compared with fine-tuning with Adam on the wav2vec2 model containing 317M training parameters, LoRA achieved similar performance by reducing the number of trainable parameters by 198 times.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bridging the Gap: A Framework for Real-World Video Deepfake Detection via Social Network Compression Emulation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A framework estimates social networks' compression settings from a few uploaded videos and reproduces those artifacts locally, so deepfake detectors can be fine-tuned without direct platform access.

Pith tools