{"work":{"id":"c59eaff8-9a9f-4650-bc74-5c7e011df1e4","openalex_id":"https://openalex.org/W4388651136","doi":"10.48550/arxiv.2311.06242","arxiv_id":"2311.06242","raw_key":null,"title":"Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks","authors":null,"authors_text":"B","year":2023,"venue":"cs.CV","abstract":"We introduce Florence-2, a novel vision foundation model with a unified, prompt-based representation for a variety of computer vision and vision-language tasks. While existing large vision models excel in transfer learning, they struggle to perform a diversity of tasks with simple instructions, a capability that implies handling the complexity of various spatial hierarchy and semantic granularity. Florence-2 was designed to take text-prompt as task instructions and generate desirable results in text forms, whether it be captioning, object detection, grounding or segmentation. This multi-task learning setup demands large-scale, high-quality annotated data. To this end, we co-developed FLD-5B that consists of 5.4 billion comprehensive visual annotations on 126 million images, using an iterative strategy of automated image annotation and model refinement. We adopted a sequence-to-sequence structure to train Florence-2 to perform versatile and comprehensive vision tasks. Extensive evaluations on numerous tasks demonstrated Florence-2 to be a strong vision foundation model contender with unprecedented zero-shot and fine-tuning capabilities.","external_url":"https://arxiv.org/abs/2311.06242","cited_by_count":14,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2311.06242","created_at":"2026-05-10T09:43:49.532421+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Florence-2: Advancing a unified representation for a variety of vision tasks (2023)","render_title":"Florence-2: Advancing a unified representation for a variety of vision tasks (2023)"},"hub":{"state":{"work_id":"c59eaff8-9a9f-4650-bc74-5c7e011df1e4","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":15,"external_cited_by_count":14,"distinct_field_count":2,"first_pith_cited_at":"2024-07-10T14:57:46+00:00","last_pith_cited_at":"2026-07-07T18:19:10+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T07:30:12.469940+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":1}],"polarity_counts":[{"context_polarity":"background","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}