{"work":{"id":"5dfe19d5-3541-4803-8fe9-3c8b9e29b281","openalex_id":"https://openalex.org/W4414682104","doi":"10.48550/arxiv.2506.15742","arxiv_id":"2506.15742","raw_key":null,"title":"FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space","authors":null,"authors_text":"Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul, Cyril Diagne","year":2025,"venue":"cs.GR","abstract":"We present evaluation results for FLUX.1 Kontext, a generative flow matching model that unifies image generation and editing. The model generates novel output views by incorporating semantic context from text and image inputs. Using a simple sequence concatenation approach, FLUX.1 Kontext handles both local editing and generative in-context tasks within a single unified architecture. Compared to current editing models that exhibit degradation in character consistency and stability across multiple turns, we observe that FLUX.1 Kontext improved preservation of objects and characters, leading to greater robustness in iterative workflows. The model achieves competitive performance with current state-of-the-art systems while delivering significantly faster generation times, enabling interactive applications and rapid prototyping workflows. To validate these improvements, we introduce KontextBench, a comprehensive benchmark with 1026 image-prompt pairs covering five task categories: local editing, global editing, character reference, style reference and text editing. Detailed evaluations show the superior performance of FLUX.1 Kontext in terms of both single-turn quality and multi-turn consistency, setting new standards for unified image processing models.","external_url":"https://arxiv.org/abs/2506.15742","cited_by_count":4,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2506.15742","created_at":"2026-05-09T06:10:41.370291+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space","render_title":"FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space"},"hub":{"state":{"work_id":"5dfe19d5-3541-4803-8fe9-3c8b9e29b281","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":237,"external_cited_by_count":4,"distinct_field_count":12,"first_pith_cited_at":"2025-08-04T11:49:20+00:00","last_pith_cited_at":"2026-07-09T07:59:29+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T04:09:25.822277+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":23},{"context_role":"baseline","n":14},{"context_role":"method","n":8},{"context_role":"other","n":1}],"polarity_counts":[{"context_polarity":"background","n":23},{"context_polarity":"baseline","n":14},{"context_polarity":"use_method","n":8},{"context_polarity":"unclear","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space","claims":[{"claim_text":"We present evaluation results for FLUX.1 Kontext, a generative flow matching model that unifies image generation and editing. The model generates novel output views by incorporating semantic context from text and image inputs. Using a simple sequence concatenation approach, FLUX.1 Kontext handles both local editing and generative in-context tasks within a single unified architecture. Compared to current editing models that exhibit degradation in character consistency and stability across multiple turns, we observe that FLUX.1 Kontext improved preservation of objects and characters, leading to ","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Nano-Banana-Pro [24] - 4.44 4.62 3.42 4.60 4.63 4.32 4.97 3.64 4.69 4.37 Seedream 4.5 [6] - 4.57 4.65 2.97 4.66 4.46 4.37 4.92 3.71 4.56 4.32 Seedream 4.0 [111] - 4.33 4.38 3.89 4.65 4.57 4.35 4.22 3.71 4.61 4.30 Nano-Banana [26] - 4.62 4.41 3.68 4.34 4.39 4.40 4.18 3.72 4.83 4.29 GPT-Image-1 [100] - 4.61 4.33 2.90 4.35 3.66 4.57 4.93 3.96 4.89 4.20 FLUX.1 Kontext [Pro] [62] - 4.25 4.15 2.35 4.56 3.57 4.26 4.57 3.68 4.63 4.00 Open-source Models Qwen-Image-Edit-2511 [139] 20B 4.54 4.57 4.13 4.70 ","claim_type":"baseline","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"training diffusion models [9,32,33], enabling continuous transport-based gener- ative learning and significantly improving sampling efficiency and visual fidelity. Recent large-scale generative models built upon the flow-matching paradigm have significantly advanced modern visual generation. Representative models in- clude image generation models such as Stable Diffusion 3.5 (SD3.5) [5], Flux [14] and Qwen-Image [41], as well as video generation models such as Wan [36], Open- Sora 2.0 [28] and K","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Chameleon: Hierarchical clustering using dynamic modeling. computer, 32(8):68-75, 1999. [11] Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. InICCV, 2023. [12] Black Forest Labs. Flux.https://github.com/black-forest-labs/flux, 2024. [13] Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul, Cyril Diagne, Tim Dockhorn, Jack English, Zion ","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"are AudioLDM2 [7], StableAudio-Open [8], and TangoFlux [4], which generate higher quality audio while addressing the generation of both general audio and music. MusicGen [9] is another high-quality open model which specializes on music generation. In this report, we propose a text-conditioned generative model specializing on instantaneous high-quality sound effect generation. Based on the multimodal FLUX-Kontext extension [10], our latent diffusion model (LDM) has been optimized from the ground ","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Equivariance regularized latent space for improved generative image modeling, 2025. [14] Taesung Kwon and Jong Chul Ye. Vision-xl: High definition video inverse problem solver using latent image diffusion models. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 10465-10474, 2025. [15] Black Forest Labs. FLUX.2: Frontier Visual Intelligence. https://bfl.ai/blog/flux-2, 2025. [16] Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"[44] Kullback, S., Leibler, R.A.: On information and sufficiency. The annals of mathemat- ical statistics22(1), 79-86 (1951) [45] Kumari, N., Zhang, B., Zhang, R., Shechtman, E., Zhu, J.Y.: Multi-concept customiza- tion of text-to-image diffusion. In: Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition. pp. 1931-1941 (2023) [46] Labs, B.F., Batifol, S., Blattmann, A., Boesel, F., Consul, S., Diagne, C., Dockhorn, T., English, J., English, Z., Esser, P ., et al.: F","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (8 contexts).","role_counts":[{"n":8,"context_role":"background"},{"n":6,"context_role":"baseline"},{"n":3,"context_role":"method"}]},"error":null,"updated_at":"2026-05-16T13:48:55.254249+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"0f6939c9-67e0-4387-8452-954f317dbcb7","orcid":null,"display_name":"Black Forest Labs"},{"id":"228c648a-898a-42f3-b70a-7c949aa5b411","orcid":null,"display_name":"Stephen Batifol"},{"id":"af305c1f-915b-4035-8a0b-28a18a2bc39f","orcid":null,"display_name":"Andreas Blattmann"},{"id":"166b8c82-1b76-48c6-8753-ac7e9061f29b","orcid":null,"display_name":"Frederic Boesel"},{"id":"cd7bd04e-019e-4e92-a083-5ba98f4fb5b2","orcid":null,"display_name":"Saksham Consul"},{"id":"56dbb868-bc37-4f40-b5e5-ece5c97bfb94","orcid":null,"display_name":"Cyril Diagne"}]},"error":null,"updated_at":"2026-05-16T13:48:56.112480+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T06:37:44.632643+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Qwen-Image Technical Report","work_id":"d06d7ecc-7579-4f89-a60b-4278a0f3c562","shared_citers":46},{"title":"Emerging Properties in Unified Multimodal Pretraining","work_id":"e0cfd82c-f5d4-44fd-b531-ec73ab0a805b","shared_citers":21},{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":21},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":21},{"title":"OmniGen2: Towards Instruction-Aligned Multimodal Generation","work_id":"d3153e5f-b6e2-4ab3-9f41-e24e24d64496","shared_citers":19},{"title":"Step1X-Edit: A Practical Framework for General Image Editing","work_id":"3392f2c8-a1cb-4d6c-8c82-2cdccffa33f9","shared_citers":18},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":17},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":16},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":15},{"title":"Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow","work_id":"a1989e1b-d66d-4533-be3a-fb9c5fd62290","shared_citers":14},{"title":"Seedream 4.0: Toward Next-generation Multimodal Image Generation","work_id":"15c839a0-48a3-4218-82b6-cac5b7f66e13","shared_citers":14},{"title":"Flow-GRPO: Training Flow Matching Models via Online RL","work_id":"bf1e8e81-ff31-401a-a5dc-d9c49df168ab","shared_citers":13},{"title":"ImgEdit: A Unified Image Editing Dataset and Benchmark","work_id":"059b5c3a-404c-4d30-a631-68c1d88a08a7","shared_citers":13},{"title":"UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation","work_id":"488a273e-95d8-46f1-87c7-2244068d00d0","shared_citers":13},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":11},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":10},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":10},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":10},{"title":"Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling","work_id":"67d9e391-26d1-459e-ab56-07e60511c886","shared_citers":10},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":10},{"title":"SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features","work_id":"50eec732-2d41-432f-9dcf-ac7fff235ea5","shared_citers":10},{"title":"BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset","work_id":"86d896d2-592f-4d9b-938e-dfeb11f9388f","shared_citers":9},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":9},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":9}],"time_series":[{"n":3,"year":2025},{"n":83,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T06:37:40.498422+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T06:37:36.278531+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space","claims":[{"claim_text":"We present evaluation results for FLUX.1 Kontext, a generative flow matching model that unifies image generation and editing. The model generates novel output views by incorporating semantic context from text and image inputs. Using a simple sequence concatenation approach, FLUX.1 Kontext handles both local editing and generative in-context tasks within a single unified architecture. Compared to current editing models that exhibit degradation in character consistency and stability across multiple turns, we observe that FLUX.1 Kontext improved preservation of objects and characters, leading to ","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Nano-Banana-Pro [24] - 4.44 4.62 3.42 4.60 4.63 4.32 4.97 3.64 4.69 4.37 Seedream 4.5 [6] - 4.57 4.65 2.97 4.66 4.46 4.37 4.92 3.71 4.56 4.32 Seedream 4.0 [111] - 4.33 4.38 3.89 4.65 4.57 4.35 4.22 3.71 4.61 4.30 Nano-Banana [26] - 4.62 4.41 3.68 4.34 4.39 4.40 4.18 3.72 4.83 4.29 GPT-Image-1 [100] - 4.61 4.33 2.90 4.35 3.66 4.57 4.93 3.96 4.89 4.20 FLUX.1 Kontext [Pro] [62] - 4.25 4.15 2.35 4.56 3.57 4.26 4.57 3.68 4.63 4.00 Open-source Models Qwen-Image-Edit-2511 [139] 20B 4.54 4.57 4.13 4.70 ","claim_type":"baseline","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"training diffusion models [9,32,33], enabling continuous transport-based gener- ative learning and significantly improving sampling efficiency and visual fidelity. Recent large-scale generative models built upon the flow-matching paradigm have significantly advanced modern visual generation. Representative models in- clude image generation models such as Stable Diffusion 3.5 (SD3.5) [5], Flux [14] and Qwen-Image [41], as well as video generation models such as Wan [36], Open- Sora 2.0 [28] and K","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Chameleon: Hierarchical clustering using dynamic modeling. computer, 32(8):68-75, 1999. [11] Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. InICCV, 2023. [12] Black Forest Labs. Flux.https://github.com/black-forest-labs/flux, 2024. [13] Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul, Cyril Diagne, Tim Dockhorn, Jack English, Zion ","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"are AudioLDM2 [7], StableAudio-Open [8], and TangoFlux [4], which generate higher quality audio while addressing the generation of both general audio and music. MusicGen [9] is another high-quality open model which specializes on music generation. In this report, we propose a text-conditioned generative model specializing on instantaneous high-quality sound effect generation. Based on the multimodal FLUX-Kontext extension [10], our latent diffusion model (LDM) has been optimized from the ground ","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Equivariance regularized latent space for improved generative image modeling, 2025. [14] Taesung Kwon and Jong Chul Ye. Vision-xl: High definition video inverse problem solver using latent image diffusion models. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 10465-10474, 2025. [15] Black Forest Labs. FLUX.2: Frontier Visual Intelligence. https://bfl.ai/blog/flux-2, 2025. [16] Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"[44] Kullback, S., Leibler, R.A.: On information and sufficiency. The annals of mathemat- ical statistics22(1), 79-86 (1951) [45] Kumari, N., Zhang, B., Zhang, R., Shechtman, E., Zhu, J.Y.: Multi-concept customiza- tion of text-to-image diffusion. In: Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition. pp. 1931-1941 (2023) [46] Labs, B.F., Batifol, S., Blattmann, A., Boesel, F., Consul, S., Diagne, C., Dockhorn, T., English, J., English, Z., Esser, P ., et al.: F","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (8 contexts).","role_counts":[{"n":8,"context_role":"background"},{"n":6,"context_role":"baseline"},{"n":3,"context_role":"method"}]},"error":null,"updated_at":"2026-05-16T13:48:56.115162+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space","claims":[{"claim_text":"We present evaluation results for FLUX.1 Kontext, a generative flow matching model that unifies image generation and editing. The model generates novel output views by incorporating semantic context from text and image inputs. Using a simple sequence concatenation approach, FLUX.1 Kontext handles both local editing and generative in-context tasks within a single unified architecture. Compared to current editing models that exhibit degradation in character consistency and stability across multiple turns, we observe that FLUX.1 Kontext improved preservation of objects and characters, leading to ","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T06:37:40.509274+00:00"}},"summary":{"title":"FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space","claims":[{"claim_text":"We present evaluation results for FLUX.1 Kontext, a generative flow matching model that unifies image generation and editing. The model generates novel output views by incorporating semantic context from text and image inputs. Using a simple sequence concatenation approach, FLUX.1 Kontext handles both local editing and generative in-context tasks within a single unified architecture. Compared to current editing models that exhibit degradation in character consistency and stability across multiple turns, we observe that FLUX.1 Kontext improved preservation of objects and characters, leading to ","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Qwen-Image Technical Report","work_id":"d06d7ecc-7579-4f89-a60b-4278a0f3c562","shared_citers":46},{"title":"Emerging Properties in Unified Multimodal Pretraining","work_id":"e0cfd82c-f5d4-44fd-b531-ec73ab0a805b","shared_citers":21},{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":21},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":21},{"title":"OmniGen2: Towards Instruction-Aligned Multimodal Generation","work_id":"d3153e5f-b6e2-4ab3-9f41-e24e24d64496","shared_citers":19},{"title":"Step1X-Edit: A Practical Framework for General Image Editing","work_id":"3392f2c8-a1cb-4d6c-8c82-2cdccffa33f9","shared_citers":18},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":17},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":16},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":15},{"title":"Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow","work_id":"a1989e1b-d66d-4533-be3a-fb9c5fd62290","shared_citers":14},{"title":"Seedream 4.0: Toward Next-generation Multimodal Image Generation","work_id":"15c839a0-48a3-4218-82b6-cac5b7f66e13","shared_citers":14},{"title":"Flow-GRPO: Training Flow Matching Models via Online RL","work_id":"bf1e8e81-ff31-401a-a5dc-d9c49df168ab","shared_citers":13},{"title":"ImgEdit: A Unified Image Editing Dataset and Benchmark","work_id":"059b5c3a-404c-4d30-a631-68c1d88a08a7","shared_citers":13},{"title":"UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation","work_id":"488a273e-95d8-46f1-87c7-2244068d00d0","shared_citers":13},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":11},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":10},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":10},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":10},{"title":"Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling","work_id":"67d9e391-26d1-459e-ab56-07e60511c886","shared_citers":10},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":10},{"title":"SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features","work_id":"50eec732-2d41-432f-9dcf-ac7fff235ea5","shared_citers":10},{"title":"BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset","work_id":"86d896d2-592f-4d9b-938e-dfeb11f9388f","shared_citers":9},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":9},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":9}],"time_series":[{"n":3,"year":2025},{"n":83,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"af305c1f-915b-4035-8a0b-28a18a2bc39f","orcid":null,"display_name":"Andreas Blattmann","source":"manual","import_confidence":0.72},{"id":"0f6939c9-67e0-4387-8452-954f317dbcb7","orcid":null,"display_name":"Black Forest Labs","source":"manual","import_confidence":0.72},{"id":"56dbb868-bc37-4f40-b5e5-ece5c97bfb94","orcid":null,"display_name":"Cyril Diagne","source":"manual","import_confidence":0.72},{"id":"166b8c82-1b76-48c6-8753-ac7e9061f29b","orcid":null,"display_name":"Frederic Boesel","source":"manual","import_confidence":0.72},{"id":"cd7bd04e-019e-4e92-a083-5ba98f4fb5b2","orcid":null,"display_name":"Saksham Consul","source":"manual","import_confidence":0.72},{"id":"228c648a-898a-42f3-b70a-7c949aa5b411","orcid":null,"display_name":"Stephen Batifol","source":"manual","import_confidence":0.72}]}}