Pith. sign in

REVIEW 11 cited by

YOLO-World: Real-Time Open-Vocabulary Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.17270 v3 pith:B6L6IBGY submitted 2024-01-30 cs.CV

classification cs.CV
keywords yolo-worlddetectionobjectopen-vocabularyachievesvision-languageyoloaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The You Only Look Once (YOLO) series of detectors have established themselves as efficient and practical tools. However, their reliance on predefined and trained object categories limits their applicability in open scenarios. Addressing this limitation, we introduce YOLO-World, an innovative approach that enhances YOLO with open-vocabulary detection capabilities through vision-language modeling and pre-training on large-scale datasets. Specifically, we propose a new Re-parameterizable Vision-Language Path Aggregation Network (RepVL-PAN) and region-text contrastive loss to facilitate the interaction between visual and linguistic information. Our method excels in detecting a wide range of objects in a zero-shot manner with high efficiency. On the challenging LVIS dataset, YOLO-World achieves 35.4 AP with 52.0 FPS on V100, which outperforms many state-of-the-art methods in terms of both accuracy and speed. Furthermore, the fine-tuned YOLO-World achieves remarkable performance on several downstream tasks, including object detection and open-vocabulary instance segmentation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Railway Artificial Intelligence Learning Benchmark (RAIL-BENCH): A Benchmark Suite for Perception in the Railway Domain

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    RAIL-BENCH is the first standardized benchmark suite for railway perception with five challenges, real-world datasets, and a novel LineAP metric for rail track detection.

  2. 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection

    cs.CV 2025-07 conditional novelty 7.0 of 10

    3D-MOOD is the first end-to-end monocular 3D object detector for open-set classes and novel scenes, achieving SOTA on Omni3D and on new open-set benchmarks.

  3. ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation

    cs.HC 2026-08 conditional novelty 6.0 of 10

    An LLM-agent artwork annotation system that combines proactive label suggestions with interaction-driven skill learning reported roughly 50% faster annotation and higher label agreement in a 12-participant study.

  4. Energy Constrained Hierarchical Underwater Monitoring via Local Multi-Agent RAG

    cs.IR 2026-07 conditional novelty 6.0 of 10

    A master–satellite edge station pairs MAX78000/02 always-on visual/acoustic sentinels with selective Jetson multimodal RAG, local species ID, and multi-agent reporting to cut energy and uplink cost.

  5. RoboAtlas: Contextual Active SLAM

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    RoboAtlas integrates frontier exploration, global semantic maps, and VLM reasoning via contextual multi-armed bandit to achieve 90.6% success rate on GOAT-Bench Val Unseen and 100% task success in real 1800m2 environments.

  6. VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors

    cs.CV 2025-10 conditional novelty 6.0 of 10

    An IoU-weighted entropy objective and image-conditioned prompt selection adapt YOLO-World and Grounding DINO at test time, improving robustness on style, weather, low-light, and corruption shifts without labels.

  7. COMPASS: Confined-space Manipulation Planning with Active Sensing Strategy

    cs.RO 2025-09 unverdicted novelty 6.0 of 10

    COMPASS is a manipulation-aware active sensing framework that raises simulated manipulation success rates by 24.25% over information-gain-only baselines in a new four-level confined-space benchmark.

  8. Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method

    cs.CV 2026-03 conditional novelty 5.0 of 10

    A progressive curriculum that trains open-vocabulary detectors on low-ambiguity, high-signal cross-modal alignments first improves robustness to visual domain shifts, with modest, test-tuned gains.

  9. MSNav: Zero-Shot Vision-and-Language Navigation with Dynamic Memory and LLM Spatial Reasoning

    cs.CV 2025-08 conditional novelty 5.0 of 10

    MSNav integrates dynamic map pruning, fine-tuned spatial reasoning (Qwen-Sp), and GPT-4o planning to improve zero-shot vision-and-language navigation on R2R and REVERIE.

  10. UrbanClipAtlas: A Visual Analytics Framework for Event and Scene Retrieval in Urban Videos

    cs.HC 2026-04 unverdicted novelty 4.0 of 10

    UrbanClipAtlas integrates RAG, taxonomy-aware extraction, and video grounding into a chat interface for retrieving and interpreting events in long urban videos from street intersections.

  11. Policy-Driven Transfer Learning in Resource-Limited Animal Monitoring

    cs.CV 2025-09 reject novelty 3.0 of 10

    Using a UCB bandit selection, the framework identifies RTDETRx as the best pre-trained model for animal detection with F1=0.718, versus 0.690 for exhaustive selection, while running fewer models per image.

Pith tools