Pith. sign in

REVIEW 6 cited by

Convolution Meets LoRA: Parameter Efficient Finetuning for Segment Anything Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.17868 v1 pith:HMFLAK6F submitted 2024-01-31 cs.CV cs.LG

Convolution Meets LoRA: Parameter Efficient Finetuning for Segment Anything Model

classification cs.CV cs.LG
keywords conv-lorasegmentationanythingdomainsimageloramodelsegment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The Segment Anything Model (SAM) stands as a foundational framework for image segmentation. While it exhibits remarkable zero-shot generalization in typical scenarios, its advantage diminishes when applied to specialized domains like medical imagery and remote sensing. To address this limitation, this paper introduces Conv-LoRA, a simple yet effective parameter-efficient fine-tuning approach. By integrating ultra-lightweight convolutional parameters into Low-Rank Adaptation (LoRA), Conv-LoRA can inject image-related inductive biases into the plain ViT encoder, further reinforcing SAM's local prior assumption. Notably, Conv-LoRA not only preserves SAM's extensive segmentation knowledge but also revives its capacity of learning high-level image semantics, which is constrained by SAM's foreground-background segmentation pretraining. Comprehensive experimentation across diverse benchmarks spanning multiple domains underscores Conv-LoRA's superiority in adapting SAM to real-world semantic segmentation tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Affordance2Action: Task-Conditioned Scene-level Affordance Grounding for Real-Time Manipulation

    cs.RO 2026-06 unverdicted novelty 7.0

    Affordance2Action introduces A2A-Bench, a manipulation-oriented benchmark for scene-level task-conditioned affordance grounding covering single- and multi-region correspondences, plus an annotation pipeline, and repor...

  2. Reclaiming Residual Knowledge: A Novel Paradigm to Low-Bit Quantization

    cs.CV 2024-08 unverdicted novelty 7.0

    CoRa reclaims quantization residuals in pre-trained ConvNets by searching low-rank adapter architectures instead of weights, matching SOTA accuracy on ImageNet in 3-4 bit settings with under 250 iterations on 1600 images.

  3. CLIP-Guided SAM: Parameter-Efficient Semantic Conditioning for Promptable Segmentation

    cs.CV 2026-05 unverdicted novelty 5.0

    CLIP-Guided SAM injects CLIP-derived features into SAM via lightweight adapters for semantic conditioning, supporting text and spatial prompts while remaining parameter-efficient and achieving competitive performance.

  4. M$^4$-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection

    cs.CV 2026-05 unverdicted novelty 5.0

    M⁴-SAM equips SAM2 with modality-aware MoE-LoRA, gated multi-level fusion, and pseudo-guided initialization to reach state-of-the-art on RGB-D video salient object detection.

  5. Generalized SAM: Efficient Fine-Tuning of SAM for Variable Input Image Sizes

    cs.CV 2024-08 unverdicted novelty 4.0

    GSAM applies random cropping to enable variable input sizes for efficient SAM fine-tuning, claiming lower compute with comparable or higher accuracy on varied datasets.

  6. Dante: An Open Source Model Pre-Training and Fine-Tuning Tool for the Dafne Federated Framework for Medical Image Segmentation

    eess.IV 2026-05 unverdicted novelty 3.0

    Dante is a new open-source backend for the Dafne ecosystem that implements configurable training from scratch, layer freezing, and channel-wise LoRA for medical image segmentation, with validation showing faster conve...