REVIEW 5 cited by
TaskCLIP: Extend Large Vision-Language Model for Task Oriented Object Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Task-oriented object detection aims to find objects suitable for accomplishing specific tasks. As a challenging task, it requires simultaneous visual data processing and reasoning under ambiguous semantics. Recent solutions are mainly all-in-one models. However, the object detection backbones are pre-trained without text supervision. Thus, to incorporate task requirements, their intricate models undergo extensive learning on a highly imbalanced and scarce dataset, resulting in capped performance, laborious training, and poor generalizability. In contrast, we propose TaskCLIP, a more natural two-stage design composed of general object detection and task-guided object selection. Particularly for the latter, we resort to the recently successful large Vision-Language Models (VLMs) as our backbone, which provides rich semantic knowledge and a uniform embedding space for images and texts. Nevertheless, the naive application of VLMs leads to sub-optimal quality, due to the misalignment between embeddings of object images and their visual attributes, which are mainly adjective phrases. To this end, we design a transformer-based aligner after the pre-trained VLMs to re-calibrate both embeddings. Finally, we employ a trainable score function to post-process the VLM matching results for object selection. Experimental results demonstrate that our TaskCLIP outperforms the state-of-the-art DETR-based model TOIST by 3.5% and only requires a single NVIDIA RTX 4090 for both training and inference.
Forward citations
Cited by 5 Pith papers
-
TRINE: A Token-Aware, Runtime-Adaptive FPGA Inference Engine for Multimodal AI
TRINE unifies ViT/CNN/GNN/NLP as DDMM/SDDMM/SpMM on a runtime mode-switchable FPGA PE array with in-stream top-k pruning and DALO scheduling, cutting latency up to 22.57× vs RTX 4090 at ~21 W.
-
Leverage Task Context for Object Affordance Ranking
The authors define task-context-conditioned object affordance ranking, release a 50,000-image benchmark, and report that their Context-embed Group Ranking model beats six saliency-ranking and multimodal-detection base...
-
Continuous GNN-based Anomaly Detection on Edge using Efficient Adaptive Knowledge Graph Learning
A GNN-based video anomaly detector adapts its knowledge graph on-device through token-embedding updates, pruning, and node creation, avoiding cloud-based graph regeneration as anomaly types change.
-
J-DDL: Surface Damage Detection and Localization System for Fighter Aircraft
A rail-based 2D/3D aircraft inspection system with an improved YOLOv8 detector (AIR-YOLO) detects damage in images and localizes it in point clouds, plus a new 8,091-image damage dataset (AIRSD).
-
Tell Me What to Track: Infusing Robust Language Guidance for Enhanced Referring Multi-Object Tracking
TellTrack improves referring multi-object tracking by adding collaborative query matching, direct query-level language infusion, and a reordered cross-modal encoder, achieving SOTA HOTA on Refer-KITTI and Refer-KITTI-V2.
Discussion (0). Continue with ORCID to comment.