REVIEW 13 cited by
YOLO-World: Real-Time Open-Vocabulary Object Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The You Only Look Once (YOLO) series of detectors have established themselves as efficient and practical tools. However, their reliance on predefined and trained object categories limits their applicability in open scenarios. Addressing this limitation, we introduce YOLO-World, an innovative approach that enhances YOLO with open-vocabulary detection capabilities through vision-language modeling and pre-training on large-scale datasets. Specifically, we propose a new Re-parameterizable Vision-Language Path Aggregation Network (RepVL-PAN) and region-text contrastive loss to facilitate the interaction between visual and linguistic information. Our method excels in detecting a wide range of objects in a zero-shot manner with high efficiency. On the challenging LVIS dataset, YOLO-World achieves 35.4 AP with 52.0 FPS on V100, which outperforms many state-of-the-art methods in terms of both accuracy and speed. Furthermore, the fine-tuned YOLO-World achieves remarkable performance on several downstream tasks, including object detection and open-vocabulary instance segmentation.
Forward citations
Cited by 13 Pith papers
-
Railway Artificial Intelligence Learning Benchmark (RAIL-BENCH): A Benchmark Suite for Perception in the Railway Domain
RAIL-BENCH is the first standardized benchmark suite for railway perception with five challenges, real-world datasets, and a novel LineAP metric for rail track detection.
-
3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
3D-MOOD is the first end-to-end monocular 3D object detector for open-set classes and novel scenes, achieving SOTA on Omni3D and on new open-set benchmarks.
-
UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation
UniMC uses tokenized instance conditions (class, box, keypoints) and a timestep-aware modulator in a DiT backbone to control multi-class human and animal image generation, trained and evaluated on the new HAIG-2.9M dataset.
-
ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation
An LLM-agent artwork annotation system that combines proactive label suggestions with interaction-driven skill learning reported roughly 50% faster annotation and higher label agreement in a 12-participant study.
-
Energy Constrained Hierarchical Underwater Monitoring via Local Multi-Agent RAG
A master–satellite edge station pairs MAX78000/02 always-on visual/acoustic sentinels with selective Jetson multimodal RAG, local species ID, and multi-agent reporting to cut energy and uplink cost.
-
RoboAtlas: Contextual Active SLAM
RoboAtlas integrates frontier exploration, global semantic maps, and VLM reasoning via contextual multi-armed bandit to achieve 90.6% success rate on GOAT-Bench Val Unseen and 100% task success in real 1800m2 environments.
-
VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors
An IoU-weighted entropy objective and image-conditioned prompt selection adapt YOLO-World and Grounding DINO at test time, improving robustness on style, weather, low-light, and corruption shifts without labels.
-
COMPASS: Confined-space Manipulation Planning with Active Sensing Strategy
COMPASS is a manipulation-aware active sensing framework that raises simulated manipulation success rates by 24.25% over information-gain-only baselines in a new four-level confined-space benchmark.
-
Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes?
Inpainting unusual objects into street scenes reveals that open-vocabulary detectors miss objects based on image location rather than object semantics.
-
Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method
A progressive curriculum that trains open-vocabulary detectors on low-ambiguity, high-signal cross-modal alignments first improves robustness to visual domain shifts, with modest, test-tuned gains.
-
MSNav: Zero-Shot Vision-and-Language Navigation with Dynamic Memory and LLM Spatial Reasoning
MSNav integrates dynamic map pruning, fine-tuned spatial reasoning (Qwen-Sp), and GPT-4o planning to improve zero-shot vision-and-language navigation on R2R and REVERIE.
-
UrbanClipAtlas: A Visual Analytics Framework for Event and Scene Retrieval in Urban Videos
UrbanClipAtlas integrates RAG, taxonomy-aware extraction, and video grounding into a chat interface for retrieving and interpreting events in long urban videos from street intersections.
-
Policy-Driven Transfer Learning in Resource-Limited Animal Monitoring
Using a UCB bandit selection, the framework identifies RTDETRx as the best pre-trained model for animal detection with F1=0.718, versus 0.690 for exhaustive selection, while running fewer models per image.
Discussion (0). Sign in to comment.