REVIEW 3 cited by
Lenna: Language Enhanced Reasoning Detection Assistant
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the fast-paced development of multimodal large language models (MLLMs), we can now converse with AI systems in natural languages to understand images. However, the reasoning power and world knowledge embedded in the large language models have been much less investigated and exploited for image perception tasks. In this paper, we propose Lenna, a language-enhanced reasoning detection assistant, which utilizes the robust multimodal feature representation of MLLMs, while preserving location information for detection. This is achieved by incorporating an additional <DET> token in the MLLM vocabulary that is free of explicit semantic context but serves as a prompt for the detector to identify the corresponding position. To evaluate the reasoning capability of Lenna, we construct a ReasonDet dataset to measure its performance on reasoning-based detection. Remarkably, Lenna demonstrates outstanding performance on ReasonDet and comes with significantly low training costs. It also incurs minimal transferring overhead when extended to other tasks. Our code and model will be available at https://git.io/Lenna.
Forward citations
Cited by 3 Pith papers
-
Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO
A recalibrated GRPO reinforcement learning method lets multimodal LLMs say 'None' for nonexistent referring expressions without sacrificing localization accuracy on objects that do exist.
-
FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing
FAS-R1 combines long-CoT supervised fine-tuning with difficulty-aware GRPO and degradation-simulated augmentation to improve multi-task face anti-spoofing and explainable rationales.
-
Language-guided Learning for Object Detection Tackling Multiple Variations in Aerial Images
LANGO adds an LLM-based visual semantic reasoner and a relation learning loss that aligns visual features with language representations, improving aerial detection AP on UAVDT and VisDrone.
Discussion (0). Continue with ORCID to comment.